Tony Wang10 min readBox Office Mojo Tags 5 of A24's 10 Best-Known Films. The Franchise Field Right Next to It Is Near-Perfect.
Two taxonomies on the same records: one matches real filmographies within a film, the other is a decorative tag. Here's the test that tells them apart.
There is a question people ask about a dataset — is this source reliable? — that turns out to be the wrong question. We went looking for a straightforward story in Box Office Mojo's data and instead found two taxonomies sitting on the same records, from the same publisher, one of which is accurate to within a single film and the other of which is decorative. Nothing on the page tells you which is which.
Here is how we found out, and the ten-minute test that would have told us sooner.
The thing we tried to write, and why it collapsed
The obvious study here writes itself: compare studios by output and by yield. The dataset makes it look easy — filter by brand, count titles, sum the grosses:
| Brand | Charted titles | Total worldwide | Median per film |
|---|---|---|---|
| Blumhouse Productions | 50 | $4.6B | $84.4M |
| Stephen King | 43 | $2.8B | $21.2M |
| Pixar | 31 | $18.8B | $539.0M |
| Illumination Entertainment | 17 | $12.2B | $654.5M |
Blumhouse ships 1.6× Pixar's output for a quarter of the money; Illumination's median is nearly eight times Blumhouse's. It is a clean, quotable finding, and it is meaningless — because those title counts are not filmographies. They are a tag, applied to some films and not others, at a rate that varies by studio.
The test: ten films you already know the answer to
For each studio, we took ten of its best-known films, confirmed each one is in the dataset, and checked whether it carries that studio's brand tag. This is the whole method. It costs a few minutes and it is the only thing that would have caught the problem.
every sampled film tagged
Us and M3GAN missing
half the sample missing
The A24 misses are not obscure. These are the films the label is known for, all present in the dataset, all with complete worldwide grosses, none carrying a brand:
| Title | Lifetime worldwide | Brand tag |
|---|---|---|
| Moonlight (2016) | $65,266,363 | none |
| The Whale (2022) | $56,772,042 | none |
| Midsommar (2019) | $47,910,822 | none |
| Ex Machina (2014) | $36,869,414 | none |
| The Lighthouse (2019) | $18,161,384 | none |
Note what the last column rules out. If these films were tagged to a co-producer or a distributor instead, the brand field would merely be ambiguous. It isn't ambiguous — it is empty.
It is not the scraper
Our first assumption was that this was our own crawl dropping tags. It isn't. We fetched Box Office Mojo's own brand pages and compared their title lists to ours, studio by studio:
| Brand | Titles in dataset | Titles on BOM's own page | Match |
|---|---|---|---|
| Pixar | 31 | 31 | exact |
| Blumhouse Productions | 50 | 50 | exact |
| Lucasfilm | 28 | 28 | exact |
| Walt Disney Animation Studios | 18 | 18 | exact |
| Illumination Entertainment | 17 | 17 | exact |
| Studio Ghibli | 17 | 17 | exact |
| A24 | 10 | 10 | exact |
That matters for anyone evaluating a scraped dataset: the gap you find is not automatically the scraper's. Here the crawl is a faithful mirror, and the sparseness is a property of the source's own editorial tagging. Checking that distinction took one request per studio.
We looked for the rule and did not find one
An uneven tag is workable if you know the rule. We tested the two obvious ones and both failed.
It is not recency. Studio Ghibli's tagged films run back to 1986, with none since 2020. Walt Disney Animation Studios, whose output starts in 1937, has no tagged film earlier than 2007. Those two point in opposite directions.
It is not box office. This is the cleanest refutation: Blumhouse's two untagged films in our sample are Us at $256.1M and M3GAN at $180.1M — among the biggest the company has ever released — while smaller Blumhouse titles are tagged. On the A24 side, Moonlight ($65.3M) is untagged and Uncut Gems ($50.0M) is tagged.
We are not going to invent a third theory. The honest statement is that the tag's coverage is not predictable from anything visible in the data, which is exactly what makes it dangerous: it fails silently, and every studio still returns a plausible-looking number.
The same records carry a taxonomy that works
Here is the part that turns this from a complaint into a method. franchise_names sits on the identical records, looks identical in a response, and is accurate. We ran the same test — ten series whose real entry counts are well documented:
| Franchise | In dataset | Real entries | Difference |
|---|---|---|---|
| James Bond | 26 | 26 | 0 |
| Friday the 13th | 12 | 12 | 0 |
| Saw | 10 | 10 | 0 |
| Rocky (incl. Creed) | 9 | 9 | 0 |
| Alien (incl. AVP) | 9 | 9 | 0 |
| Scream | 7 | 7 | 0 |
| Terminator | 6 | 6 | 0 |
| Star Wars | 13 | 12 | +1 |
| Halloween | 12 | 13 | −1 |
| Jurassic Park | 6 | 7 | −1 |
One publisher. One dataset. One record. Two tags, and a reliability gap you cannot see from the response.
The test, generalised
Nothing above required domain expertise in film. It required picking entities whose answer we already knew and counting. That generalises to any dataset field you are about to build on:
- Choose ten entities you can verify independently — not ten at random. Well-known ones, where a missing value is obvious.
- Confirm the records exist first. A missing tag and a missing record are different bugs with different fixes, and conflating them sends you after the wrong one.
- Check whether a gap is empty or contested. "No brand" and "a different brand" mean opposite things about the field.
- Test the gap against the two obvious rules — time and size. If neither explains it, treat the field as unpredictable rather than inventing a third rule.
- Go to the source before blaming the pipeline. One request per entity settled it here, and it moved the finding from "our crawl is lossy" to "the source's taxonomy is sparse" — a completely different conclusion.
Three more traps in the same dataset
Found while testing the above, each worth knowing before you query it:
lifetime_yearis missing on precisely the films you'd need it for. Coverage is about two-thirds overall, but it is not random: Marvel Cinematic Universe entries have it on 38 of 38, Star Wars 12 of 13 — and Friday the 13th on 0 of 12, Halloween 2 of 12, James Bond 10 of 26. The old films are the ones without years, so any inflation adjustment or era comparison is impossible in the direction you'd want it.- Two documented facets are not filters.
in_lifetime_top_1000_wwandyears_activecan be faceted but are silently ignored by the search endpoint, which returns the full 21,973 rows with a 200 — the same thing it does for a parameter you invented. Our first attempt at an inflation-immune metric reported that 100% of all fifty franchises were in the all-time worldwide top 1000; the facet says only 1,004 titles qualify. Always confirm a filter actually reduced the total. - The genre taxonomy is a keyword list, not genres.
Supernaturalis applied to Avengers: Endgame, Barbie, Frozen and Aladdin. There is no plainHorrorvalue at all — onlySlasher(102 titles) andHorror Comedy(127). Genre facets also cap at 50 buckets, withDocumentaryandForeign Languageboth truncated at exactly 2000.
What the data does support
Stripped of the fields that don't hold up, the population itself is solid and worth citing: of 21,973 charted titles, 81% never cleared $50M in lifetime worldwide gross, and 65 — 0.3% — cleared $1B. 1,004 sit in the all-time worldwide top 1000 by construction. Franchise-level analysis is on firm ground; studio-level analysis, on this field, is not.
Charted titles, lifetime and yearly grosses, release groups, market breakdowns and franchise tags are queryable at /datasets/boxofficemojo over one REST API. If you are about to build on any field in any dataset — this one included — spend the ten minutes on the ten-entity test first.
Related reading
Frequently asked questions
Is Box Office Mojo's brand field a reliable studio filmography?
No. Of ten well-known A24 films, all present in the dataset with full worldwide grosses, Box Office Mojo tags five with the A24 brand — Moonlight, The Whale, Midsommar, Ex Machina and The Lighthouse carry no brand at all. Coverage varies by studio: a ten-film sample is tagged 9 of 9 for Pixar, 8 of 10 for Blumhouse and 5 of 10 for A24. Any study comparing studio output or per-film yield using this field is measuring tagging coverage rather than filmographies.
Are the missing brand tags a scraping problem?
No. We fetched Box Office Mojo's own /brand/ pages and compared their title lists to the dataset for seven studios: all seven match exactly — Pixar 31, Blumhouse 50, Lucasfilm 28, Disney Animation 18, Illumination 17, Studio Ghibli 17, A24 10. Box Office Mojo's own A24 brand page lists exactly ten films. The crawl is a faithful mirror; the sparseness is in the source's editorial tagging.
Which films does Box Office Mojo choose to tag with a brand?
We could not find a rule, and we tested the two obvious ones. It is not recency: Studio Ghibli's tagged films run back to 1986 with none since 2020, while Disney Animation, whose output starts in 1937, has no tagged film before 2007. It is not box office: Blumhouse's Us ($256.1M) and M3GAN ($180.1M), among the biggest films it has released, are both untagged while smaller Blumhouse titles are tagged. The coverage is not predictable from anything visible in the data.
Is Box Office Mojo's franchise data trustworthy?
Yes, on the evidence — and it sits on the same records as the brand field. Tested against ten series with well-documented entry counts, the franchise tag was exact on seven and off by one on three: James Bond 26, Friday the 13th 12, Saw 10, Rocky 9, Alien 9, Scream 7, Terminator 6, with Star Wars +1, Halloween −1 and Jurassic Park −1. Mean absolute error 0.3 films. Two fields from one publisher on one record can have opposite reliability.
How do you test whether a dataset field is trustworthy?
Pick ten entities whose answer you already know, and count. Confirm the records themselves exist first, since a missing tag and a missing record are different bugs. Check whether a gap is empty or contested — 'no value' and 'a different value' mean opposite things. Test the gap against time and size, and if neither explains it, treat the field as unpredictable rather than inventing a rule. And go to the source before blaming your pipeline: one request per entity moved this finding from 'our crawl is lossy' to 'the source's taxonomy is sparse'.
What Box Office Mojo figures are safe to cite?
The population and the franchise tags. Of 21,973 charted theatrical titles, 81% never cleared $50M lifetime worldwide and 65 — 0.3% — cleared $1B. Franchise-level analysis is on firm ground. Avoid lifetime_year for era comparisons, since it is missing on exactly the older films you would need (Friday the 13th: 0 of 12, Halloween 2 of 12, James Bond 10 of 26), and note that in_lifetime_top_1000_ww and years_active are facets the search endpoint silently ignores as filters.