Historical Quality-Weighted Share of Voice
How Muse Parlor turns a century of digitized magazines into a structured, quality-weighted measure of how beauty brands appeared in the cultural record. This document is written to be reproduced, challenged, and extended.
AbstractWhat this measures
Muse Parlor reconstructs the share of voice of named beauty houses from the digitized magazine record, decade by decade, in both editorial and advertising, back to the early twentieth century. Presence is not counted as a raw tally. Each appearance is scored for quality on the same engine as the live contemporary model, then aggregated into a quality-weighted share of voice per period.
The method is deliberately descriptive. It reports how brands appeared in the record. It does not estimate causal effect on sales or demand, because the instruments that would identify such effects, in particular branded search, do not exist before roughly 2004. Every figure is falsifiable against a named page in a named issue.
The one-sentence version
A vision model reads the actual page scans to find brand appearances that plain text search cannot see, a second pass verifies each against the printed word and an explicit entity spec, each verified appearance is scored for placement quality, and the scores are aggregated into a per-period share of voice.
Section 1Corpus and rights
The Phase 0 corpus is drawn entirely from open, digitized periodical archives reached through official search and image APIs. No collection is crawled, and no page image is redistributed. Each sampled page carries a source_id that resolves to a rights ledger recording the host, collection, access method, and license position.
| Publication | Host | Type | Window sampled |
|---|---|---|---|
| Photoplay | Internet Archive, Media History Digital Library | Fan magazine | 1914–1950 |
| Ladies' Home Journal | Internet Archive | Women's service | 1920–1930 |
| The Delineator | Internet Archive | Women's service | 1920–1930 |
Rights basis: the sampled volumes are public domain or reached under a non-consumptive research posture. For any HathiTrust-only coverage in later phases, the HathiTrust Extracted Features 2.5 dataset is used, which is released under CC-BY 4.0 and supports quantitative extraction without redistribution of page images. Provenance is a column on every row, not a footnote.
Section 2Sampling frame
Because a full census of every page is neither necessary nor economical, Muse Parlor draws a stratified page-level sample. The universe is enumerated from the host's search API by title and year. Pages are then drawn to spread evenly across the available years, with issues rotated so no single year or issue dominates, and the first four and last two leaves of each volume skipped to avoid covers and blanks.
The Phase 0 harvest reported here is 3,000 pages: 1,800 Photoplay across 1914 to 1950, 800 Ladies' Home Journal, and 400 Delineator, spanning 36 distinct years. The sampler is seeded (PHASE0_SEED) so any frame is reproducible exactly.
The empirical base rate of a real, tracked-brand mention is about 3.1 percent of pages. This low rate is the central sampling fact: to reach a per-brand count large enough to estimate precision reliably, the frame must be in the thousands of pages, not the hundreds.
Section 3Extraction, a two-stage vision method
The core technical claim of Muse Parlor is that optical character recognition is structurally insufficient for this task, and that a vision model reading the page image is required.
Why OCR fails
Beauty brand names in this era are rendered as logotypes and stylized display type. A hand-validation on the Phase 0 sample found the ceiling recall of page OCR, meaning the fraction of true mentions whose brand token appears anywhere in the OCR text at all, was about 59 percent overall, 31 percent for Coty, and 0 percent for Palmolive. The names are on the page as images, not as text. No matcher on top of OCR can recover what OCR never captured.
Stage one, a high-recall finder
Each sampled page image is read at 1,400 pixels wide by a vision model, which reports every tracked brand it can see, as text or as a logo, with a channel label (advertising or editorial), a confidence, and a short evidence note. This stage is tuned for recall. It over-generates, flagging brands from company names, product categories, and loose association.
Stage two, a grounded verifier
Every stage-one candidate is re-shown the page and required to quote the exact printed text that names the brand, and to reconcile that text against the brand's identity and an explicit list of confusables. A candidate is kept only if the quoted printed name genuinely is the brand or one of its documented lines. Candidates the model self-rejects (confidence below 0.35) are never emitted.
where Stage2 requires a printed quote q(p,b) that resolves to b under the entity spec (Section 4)
On the validation set this two-stage design rejected every adversarial trap, for example a "COLLINS" logo and a "KOTEX" logo misread as Coty, "Luxor" tagged as Lux, and film-star copy inferred as Lux, while keeping genuine, quotable advertisements. The measured effect is a large gain in precision at no measured cost in recall.
Section 4Entity resolution
The disambiguation layer is where the real work sits, and it is governed by a single rule.
Governing rule
Track the mark, not the corporate parent. Count the tracked brand name, its spelling and OCR variants, and the product lines advertised under that name. Do not count a separate mark even under the same parent, and do not count a lookalike brand.
Each of the eight seed houses has a sourced definition listing the canonical name, the sub-lines that count, and the confusables that explicitly do not. Two worked cases show why this matters.
Luxor is not Lux
Luxor began inside Armour and Company, became Luxor Ltd around 1919, and was acquired by Lever Brothers only in 1948. It later shared a corporate parent with Lux, yet it is a different mark and is excluded. The verifier enforces this: it identifies "Luxor" on the page and declines to count it as Lux.
Coty and the perfume field
Coty is a short name in a category dense with French perfume houses, and it is the hardest case. In the Phase 0 harvest the model initially attributed Djer-Kiss (Kerkoff), Vivatone (Daggett and Ramsdell), and Cashmere Bouquet (a Colgate face powder) to Coty, along with a fashion color "chypre" and a radio-skit "Station STYX". The spec now names those houses as confusables and restricts Coty to its literal name plus its distinctively owned lines (L'Origan, Emeraude, L'Aimant, Air-Spun). Ambiguous perfume words count only with the printed word Coty beside them.
| House | Counts | Does not count |
|---|---|---|
| Lux | Lux Toilet Soap | Luxor, Lux flakes and Lux for dishes (laundry), de luxe |
| Coty | Coty, L'Origan, Emeraude, L'Aimant, Air-Spun | Djer-Kiss, Vivatone, Cashmere Bouquet, generic "chypre", "Styx" as a place |
| Maybelline | Maybelline, early Maybell Laboratories | "Mabel" as a name |
| Elizabeth Arden | Elizabeth Arden, Ardena, Blue Grass, Venetian | "Arden" as a place or family name |
| Helena Rubinstein | Helena Rubinstein, the Valaze line | "Rubinstein" the pianist |
Definitions are sourced to brand histories, chiefly Cosmetics and Skin, and are versioned with the corpus. The full spec covers all eight houses.
Section 5The Placement Quality Score
A mention is not a unit of voice. A full-page Lux advertisement and a passing line in a gossip column are both "presence," but they are not equal. Muse Parlor scores each verified placement with a Placement Quality Score, mirroring the live contemporary model exactly.
The attribute scores are read from the page by a vision pass over each mention, on fixed 0-to-1 rubric anchors. The weights are inherited from the live model, with one adaptation described below.
| Attribute | Rubric anchors (0 to 1) | Weight |
|---|---|---|
| prominence | 1.0 full page or dominant, 0.6 half or quarter, 0.3 small block, 0.1 incidental | 29.4% |
| message | 1.0 brand is the subject, 0.5 featured among others, 0.0 incidental | 23.5% |
| visual | 1.0 shown with logo, product, or illustration, 0.0 text only | 11.8% |
| sentiment | 1.0 positive or aspirational, 0.6 neutral, 0.2 negative | 11.8% |
| attribution | 1.0 editorial (earned), 0.5 advertising (paid) | 11.8% |
| exclusivity | 1.0 brand alone or dominant, 0.4 shares the page | 11.8% |
Two historical adaptations
- The live model includes a hyperlink attribute, weighted 15 percent, which cannot exist before the web. It is dropped, and the remaining six weights are renormalized to sum to one. Concretely, each surviving weight is divided by 0.85, which is why prominence moves from 25.0 to 29.4 percent.
- Attribution is taken directly from the channel already captured at extraction rather than scored again, so editorial earned coverage scores above paid advertising.
Outlet authority, circulation-indexed
Authority weights the outlet by reach. It is indexed to representative period circulation with a log-compressed function, anchored so the smallest title is 1.0 and the largest lands near the top of the live model's authority band. Log compression reflects the diminishing marginal reach of raw circulation, a title with ten times the copies is not ten times the authority.
| Outlet | Circulation, monthly | Authority |
|---|---|---|
| Ladies' Home Journal | ~2,000,000 | 1.70 |
| The Delineator | ~1,000,000 | 1.44 |
| Photoplay | ~300,000 | 1.00 |
Figures are representative over the sampled window, Ladies' Home Journal about two million by the late 1910s, The Delineator about one million at its early-1920s peak, Photoplay 204,434 in 1918 and hundreds of thousands through the 1920s and 1930s, sourced to Britannica, Wikipedia, and Encyclopedia.com with the Media History Digital Library. Within a single-title series authority cancels, so it shapes cross-title comparison and absolute scores, not the constant-corpus series. A time-varying, year-level circulation index is a Phase 1 refinement.
Section 6Quality-weighted share of voice
Within a period, a house's share of voice is its share of the total placement quality of the competitive set.
The constant-corpus correction
A naive cross-title series is confounded, because the titles in the sample change across time. In this frame, Ladies' Home Journal and Delineator sit only in the 1920s window, so houses that advertised chiefly in the women's service titles appear concentrated there partly for that reason. The corrected view holds the corpus constant, computing the series within a single title (Photoplay, 1917 to 1950) so that composition is fixed and only brand behavior varies. Cross-title series are reported only with a per-title normalization.
Section 7Validation, the Phase 0 gates
The method is gated on three falsifiable criteria before any series is trusted. Ground truth is established by reading the scans directly, not by a hand-coding pass seeded from the model, which would be circular.
C1, extraction fidelity
Precision and recall of the extractor against adjudicated truth, per brand where the count reaches ten.
gate: precision ≥ 0.85 and recall ≥ 0.75
Pass After the Coty tightening, all eight brands adjudicated clean, and the four brands at count ten or more, Lux, Maybelline, Pond's, and Coty, clear the gate. A blind check of engine-empty pages found no missed mentions, a preliminary recall result to be widened.
C2, coverage sufficiency
Gate: at least five of eight brands present, at least sixty placements, at least fifteen distinct years. Pass Eight of eight brands, 88 placements, 26 distinct years across 1917 to 1950.
C3, discrimination
The series must separate brands beyond noise. Gate: between-brand variance in period share exceeds within-brand temporal variance, and the top-three rank order changes at least once. Pass On the constant corpus, between-brand variance 74 exceeds within-brand temporal variance 50 on the well-sampled decades, and the top three reorder every decade.
Section 8Limitations and posture
- Descriptive only. With no branded-search series before roughly 2004, no causal coefficient is estimated for the historical period. Presence is reported, effect is not claimed.
- Composition confound. Cross-title decade series reflect which titles were sampled as well as brand behavior. The constant-corpus view is the honest default.
- Authority is circulation-indexed but static. Outlet weights derive from representative period circulation; a time-varying, year-level circulation index is a Phase 1 refinement.
- Recall bound is preliminary. The empty-page recall check is directional and will be widened to a powered blind audit.
- Digitization and survivorship bias. The record is what was scanned and what survived, which is not a random sample of what was published.
Section 9Reproducibility
Every result here is regenerable from a seed and a small set of parameters.
| Parameter | Value |
|---|---|
| Sampler seed | PHASE0_SEED = 7 |
| Harvest | 3,000 pages, 1917–1950, 36 years |
| Extraction image width | 1,400 px (IIIF), 8-way sharded |
| Emit threshold | confidence ≥ 0.35 |
| PQS weights | live weights, link dropped, renormalized |
| Authority | circulation-indexed, LHJ 1.70, Delineator 1.44, Photoplay 1.00 |
| Approx. compute cost | ~$9 for the full harvest |
Artifacts include the sampling frame, the per-mention extraction with its printed-quote evidence, the per-mention PQS with each attribute score, and the per-period share series. Any figure can be traced back to a specific leaf of a specific issue.
Section 10For researchers
This methodology is built to be extended, and Muse Parlor intends to host research that does so. Open directions include:
- Circulation-indexed authority. Replace the provisional outlet weights with weights derived from period circulation and advertising-rate data.
- Inter-coder reliability. A formal study of agreement between the vision extractor, independent vision models, and human coders, reporting kappa on a shared blind set.
- OCR versus vision benchmark. A publishable quantification of the OCR recall ceiling for display-type brand names across periods and titles.
- Corpus expansion. Extending the frame across more titles and back into the 1890s, with per-title normalization of the cross-title series.
- Bridging to the causal era. Designs that connect the descriptive historical series to the post-2004 period where branded search enables identification.
Bring a research question to the archive
Muse Parlor works with graduate programs and researchers to design, run, and host studies on this corpus and method. Datasets, parameters, and provenance are shareable under agreement.
Propose a collaborationReferencesSources
- Brand histories, chiefly Cosmetics and Skin, cosmeticsandskin.com, company pages for Coty, Elizabeth Arden, Helena Rubinstein, Max Factor, Lux, Pond's, Palmolive, Maybelline.
- Luxor history, Collecting Vintage Compacts.
- Corpus, Internet Archive and the Media History Digital Library (Photoplay), Internet Archive (Ladies' Home Journal, The Delineator).
- Circulation figures, Britannica and Wikipedia (Ladies' Home Journal, The Delineator), Encyclopedia.com and the Media History Digital Library (Photoplay).
- HathiTrust Research Center, Extracted Features 2.5, released CC-BY 4.0.
- The live contemporary engine, MML Share of Voice and Answer Engine Optimization methodology, from which the Placement Quality Score and quality-weighted share of voice are inherited.
