Muse ParlorMuse Parlor
Muse Parlor Methodology

Historical Quality-Weighted Share of Voice

How Muse Parlor turns a century of digitized magazines into a structured, quality-weighted measure of how beauty brands appeared in the cultural record. This document is written to be reproduced, challenged, and extended.

Version Phase 0Posture DescriptiveA Black Dome Corp product

AbstractWhat this measures

Muse Parlor reconstructs the share of voice of named beauty houses from the digitized magazine record, decade by decade, in both editorial and advertising, back to the early twentieth century. Presence is not counted as a raw tally. Each appearance is scored for quality on the same engine as the live contemporary model, then aggregated into a quality-weighted share of voice per period.

The method is deliberately descriptive. It reports how brands appeared in the record. It does not estimate causal effect on sales or demand, because the instruments that would identify such effects, in particular branded search, do not exist before roughly 2004. Every figure is falsifiable against a named page in a named issue.

The one-sentence version

A vision model reads the actual page scans to find brand appearances that plain text search cannot see, a second pass verifies each against the printed word and an explicit entity spec, each verified appearance is scored for placement quality, and the scores are aggregated into a per-period share of voice.

Section 1Corpus and rights

The Phase 0 corpus is drawn entirely from open, digitized periodical archives reached through official search and image APIs. No collection is crawled, and no page image is redistributed. Each sampled page carries a source_id that resolves to a rights ledger recording the host, collection, access method, and license position.

Phase 0 corpus
PublicationHostTypeWindow sampled
PhotoplayInternet Archive, Media History Digital LibraryFan magazine1914–1950
Ladies' Home JournalInternet ArchiveWomen's service1920–1930
The DelineatorInternet ArchiveWomen's service1920–1930

Rights basis: the sampled volumes are public domain or reached under a non-consumptive research posture. For any HathiTrust-only coverage in later phases, the HathiTrust Extracted Features 2.5 dataset is used, which is released under CC-BY 4.0 and supports quantitative extraction without redistribution of page images. Provenance is a column on every row, not a footnote.

Section 2Sampling frame

Because a full census of every page is neither necessary nor economical, Muse Parlor draws a stratified page-level sample. The universe is enumerated from the host's search API by title and year. Pages are then drawn to spread evenly across the available years, with issues rotated so no single year or issue dominates, and the first four and last two leaves of each volume skipped to avoid covers and blanks.

The Phase 0 harvest reported here is 3,000 pages: 1,800 Photoplay across 1914 to 1950, 800 Ladies' Home Journal, and 400 Delineator, spanning 36 distinct years. The sampler is seeded (PHASE0_SEED) so any frame is reproducible exactly.

The empirical base rate of a real, tracked-brand mention is about 3.1 percent of pages. This low rate is the central sampling fact: to reach a per-brand count large enough to estimate precision reliably, the frame must be in the thousands of pages, not the hundreds.

Section 3Extraction, a two-stage vision method

The core technical claim of Muse Parlor is that optical character recognition is structurally insufficient for this task, and that a vision model reading the page image is required.

Why OCR fails

Beauty brand names in this era are rendered as logotypes and stylized display type. A hand-validation on the Phase 0 sample found the ceiling recall of page OCR, meaning the fraction of true mentions whose brand token appears anywhere in the OCR text at all, was about 59 percent overall, 31 percent for Coty, and 0 percent for Palmolive. The names are on the page as images, not as text. No matcher on top of OCR can recover what OCR never captured.

Stage one, a high-recall finder

Each sampled page image is read at 1,400 pixels wide by a vision model, which reports every tracked brand it can see, as text or as a logo, with a channel label (advertising or editorial), a confidence, and a short evidence note. This stage is tuned for recall. It over-generates, flagging brands from company names, product categories, and loose association.

Stage two, a grounded verifier

Every stage-one candidate is re-shown the page and required to quote the exact printed text that names the brand, and to reconcile that text against the brand's identity and an explicit list of confusables. A candidate is kept only if the quoted printed name genuinely is the brand or one of its documented lines. Candidates the model self-rejects (confidence below 0.35) are never emitted.

Extraction, per page p and brand b keep(p,b) = Stage1(p,b) ∧ Stage2(p,b) ∧ conf(p,b) ≥ 0.35
where Stage2 requires a printed quote q(p,b) that resolves to b under the entity spec (Section 4)

On the validation set this two-stage design rejected every adversarial trap, for example a "COLLINS" logo and a "KOTEX" logo misread as Coty, "Luxor" tagged as Lux, and film-star copy inferred as Lux, while keeping genuine, quotable advertisements. The measured effect is a large gain in precision at no measured cost in recall.

Section 4Entity resolution

The disambiguation layer is where the real work sits, and it is governed by a single rule.

Governing rule

Track the mark, not the corporate parent. Count the tracked brand name, its spelling and OCR variants, and the product lines advertised under that name. Do not count a separate mark even under the same parent, and do not count a lookalike brand.

Each of the eight seed houses has a sourced definition listing the canonical name, the sub-lines that count, and the confusables that explicitly do not. Two worked cases show why this matters.

Luxor is not Lux

Luxor began inside Armour and Company, became Luxor Ltd around 1919, and was acquired by Lever Brothers only in 1948. It later shared a corporate parent with Lux, yet it is a different mark and is excluded. The verifier enforces this: it identifies "Luxor" on the page and declines to count it as Lux.

Coty and the perfume field

Coty is a short name in a category dense with French perfume houses, and it is the hardest case. In the Phase 0 harvest the model initially attributed Djer-Kiss (Kerkoff), Vivatone (Daggett and Ramsdell), and Cashmere Bouquet (a Colgate face powder) to Coty, along with a fashion color "chypre" and a radio-skit "Station STYX". The spec now names those houses as confusables and restricts Coty to its literal name plus its distinctively owned lines (L'Origan, Emeraude, L'Aimant, Air-Spun). Ambiguous perfume words count only with the printed word Coty beside them.

Entity spec, abridged
HouseCountsDoes not count
LuxLux Toilet SoapLuxor, Lux flakes and Lux for dishes (laundry), de luxe
CotyCoty, L'Origan, Emeraude, L'Aimant, Air-SpunDjer-Kiss, Vivatone, Cashmere Bouquet, generic "chypre", "Styx" as a place
MaybellineMaybelline, early Maybell Laboratories"Mabel" as a name
Elizabeth ArdenElizabeth Arden, Ardena, Blue Grass, Venetian"Arden" as a place or family name
Helena RubinsteinHelena Rubinstein, the Valaze line"Rubinstein" the pianist

Definitions are sourced to brand histories, chiefly Cosmetics and Skin, and are versioned with the corpus. The full spec covers all eight houses.

Section 5The Placement Quality Score

A mention is not a unit of voice. A full-page Lux advertisement and a passing line in a gossip column are both "presence," but they are not equal. Muse Parlor scores each verified placement with a Placement Quality Score, mirroring the live contemporary model exactly.

Placement Quality Score PQS(m) = authority(outletm) × Σa scorea(m) · wa

The attribute scores are read from the page by a vision pass over each mention, on fixed 0-to-1 rubric anchors. The weights are inherited from the live model, with one adaptation described below.

Attributes, rubric anchors, and weights
AttributeRubric anchors (0 to 1)Weight
prominence1.0 full page or dominant, 0.6 half or quarter, 0.3 small block, 0.1 incidental29.4%
message1.0 brand is the subject, 0.5 featured among others, 0.0 incidental23.5%
visual1.0 shown with logo, product, or illustration, 0.0 text only11.8%
sentiment1.0 positive or aspirational, 0.6 neutral, 0.2 negative11.8%
attribution1.0 editorial (earned), 0.5 advertising (paid)11.8%
exclusivity1.0 brand alone or dominant, 0.4 shares the page11.8%

Two historical adaptations

  1. The live model includes a hyperlink attribute, weighted 15 percent, which cannot exist before the web. It is dropped, and the remaining six weights are renormalized to sum to one. Concretely, each surviving weight is divided by 0.85, which is why prominence moves from 25.0 to 29.4 percent.
  2. Attribution is taken directly from the channel already captured at extraction rather than scored again, so editorial earned coverage scores above paid advertising.

Outlet authority, circulation-indexed

Authority weights the outlet by reach. It is indexed to representative period circulation with a log-compressed function, anchored so the smallest title is 1.0 and the largest lands near the top of the live model's authority band. Log compression reflects the diminishing marginal reach of raw circulation, a title with ten times the copies is not ten times the authority.

Outlet authority authority(o) = 1 + 0.85 · log10( circulation(o) ÷ circulationmin )
Circulation-indexed authority
OutletCirculation, monthlyAuthority
Ladies' Home Journal~2,000,0001.70
The Delineator~1,000,0001.44
Photoplay~300,0001.00

Figures are representative over the sampled window, Ladies' Home Journal about two million by the late 1910s, The Delineator about one million at its early-1920s peak, Photoplay 204,434 in 1918 and hundreds of thousands through the 1920s and 1930s, sourced to Britannica, Wikipedia, and Encyclopedia.com with the Media History Digital Library. Within a single-title series authority cancels, so it shapes cross-title comparison and absolute scores, not the constant-corpus series. A time-varying, year-level circulation index is a Phase 1 refinement.

Section 6Quality-weighted share of voice

Within a period, a house's share of voice is its share of the total placement quality of the competitive set.

Quality-weighted share of voice, house b, period t QWSoV(b,t) = [ Σm ∈ (b,t) PQS(m) ] ÷ [ Σb' Σm ∈ (b',t) PQS(m) ] × 100

The constant-corpus correction

A naive cross-title series is confounded, because the titles in the sample change across time. In this frame, Ladies' Home Journal and Delineator sit only in the 1920s window, so houses that advertised chiefly in the women's service titles appear concentrated there partly for that reason. The corrected view holds the corpus constant, computing the series within a single title (Photoplay, 1917 to 1950) so that composition is fixed and only brand behavior varies. Cross-title series are reported only with a per-title normalization.

Section 7Validation, the Phase 0 gates

The method is gated on three falsifiable criteria before any series is trusted. Ground truth is established by reading the scans directly, not by a hand-coding pass seeded from the model, which would be circular.

C1, extraction fidelity

Precision and recall of the extractor against adjudicated truth, per brand where the count reaches ten.

precision = TP ÷ (TP + FP)  recall = TP ÷ (TP + FN)
gate: precision ≥ 0.85 and recall ≥ 0.75

Pass After the Coty tightening, all eight brands adjudicated clean, and the four brands at count ten or more, Lux, Maybelline, Pond's, and Coty, clear the gate. A blind check of engine-empty pages found no missed mentions, a preliminary recall result to be widened.

C2, coverage sufficiency

Gate: at least five of eight brands present, at least sixty placements, at least fifteen distinct years. Pass Eight of eight brands, 88 placements, 26 distinct years across 1917 to 1950.

C3, discrimination

The series must separate brands beyond noise. Gate: between-brand variance in period share exceeds within-brand temporal variance, and the top-three rank order changes at least once. Pass On the constant corpus, between-brand variance 74 exceeds within-brand temporal variance 50 on the well-sampled decades, and the top three reorder every decade.

Section 8Limitations and posture

  • Descriptive only. With no branded-search series before roughly 2004, no causal coefficient is estimated for the historical period. Presence is reported, effect is not claimed.
  • Composition confound. Cross-title decade series reflect which titles were sampled as well as brand behavior. The constant-corpus view is the honest default.
  • Authority is circulation-indexed but static. Outlet weights derive from representative period circulation; a time-varying, year-level circulation index is a Phase 1 refinement.
  • Recall bound is preliminary. The empty-page recall check is directional and will be widened to a powered blind audit.
  • Digitization and survivorship bias. The record is what was scanned and what survived, which is not a random sample of what was published.

Section 9Reproducibility

Every result here is regenerable from a seed and a small set of parameters.

Phase 0 parameters
ParameterValue
Sampler seedPHASE0_SEED = 7
Harvest3,000 pages, 1917–1950, 36 years
Extraction image width1,400 px (IIIF), 8-way sharded
Emit thresholdconfidence ≥ 0.35
PQS weightslive weights, link dropped, renormalized
Authoritycirculation-indexed, LHJ 1.70, Delineator 1.44, Photoplay 1.00
Approx. compute cost~$9 for the full harvest

Artifacts include the sampling frame, the per-mention extraction with its printed-quote evidence, the per-mention PQS with each attribute score, and the per-period share series. Any figure can be traced back to a specific leaf of a specific issue.

Section 10For researchers

This methodology is built to be extended, and Muse Parlor intends to host research that does so. Open directions include:

  • Circulation-indexed authority. Replace the provisional outlet weights with weights derived from period circulation and advertising-rate data.
  • Inter-coder reliability. A formal study of agreement between the vision extractor, independent vision models, and human coders, reporting kappa on a shared blind set.
  • OCR versus vision benchmark. A publishable quantification of the OCR recall ceiling for display-type brand names across periods and titles.
  • Corpus expansion. Extending the frame across more titles and back into the 1890s, with per-title normalization of the cross-title series.
  • Bridging to the causal era. Designs that connect the descriptive historical series to the post-2004 period where branded search enables identification.
Collaboration

Bring a research question to the archive

Muse Parlor works with graduate programs and researchers to design, run, and host studies on this corpus and method. Datasets, parameters, and provenance are shareable under agreement.

Propose a collaboration

ReferencesSources

  • Brand histories, chiefly Cosmetics and Skin, cosmeticsandskin.com, company pages for Coty, Elizabeth Arden, Helena Rubinstein, Max Factor, Lux, Pond's, Palmolive, Maybelline.
  • Luxor history, Collecting Vintage Compacts.
  • Corpus, Internet Archive and the Media History Digital Library (Photoplay), Internet Archive (Ladies' Home Journal, The Delineator).
  • Circulation figures, Britannica and Wikipedia (Ladies' Home Journal, The Delineator), Encyclopedia.com and the Media History Digital Library (Photoplay).
  • HathiTrust Research Center, Extracted Features 2.5, released CC-BY 4.0.
  • The live contemporary engine, MML Share of Voice and Answer Engine Optimization methodology, from which the Placement Quality Score and quality-weighted share of voice are inherited.