Tony Wang9 min readX's Everyday Accounts List Their Location More Than Its Notable Ones Do
A 2,462-account X audit found 63% fill in location — Common Crawl accounts (65%) beat curated, notable ones (51%), flipping this series' usual pattern.
Three studies in this series have found the same pattern: the more reach an X account has, the more complete its profile. Verified accounts fill in a bio more than unverified ones. Curated, notable accounts fill in a bio more than accounts our newer Common Crawl channel discovered through ordinary web mentions. We went looking for whether that pattern holds for one more field — location — and expected a fourth confirmation. Instead we found the one place it flips.
Location isn't in the bulk data — we had to ask account by account
Every other stat in this series came from our public search API, which returns bulk facets and filtered counts instantly, no per-record fetching required. Location isn't one of those fields — the bulk search and item records carry username, bio, follower counts, verification, and source tier, but not location. It only appears in the response from a live, per-account profile lookup.
So this study is built differently from its predecessors: a stratified sample of 2,482 usernames, drawn across all eight major discovery tiers and six follower bands (to avoid oversampling mega-accounts), each looked up individually. 2,462 resolved (99.2% — 16 were private or suspended, 4 timed out). Location was empty on 901 of those and filled in on 1,561 — a 63.4% fill rate overall.
The one signal where "ordinary" beats "notable"
| Discovery tier | Sample size | Location fill rate |
|---|---|---|
| GitHub users | 583 | 76.8% |
| X search (authenticated) | 202 | 70.3% |
| Journalists | 148 | 66.9% |
| Common Crawl | 599 | 64.8% |
| TrustMRR startups | 92 | 63.0% |
| TikTok creators | 100 | 52.0% |
| Wikidata (curated/notable) | 596 | 51.2% |
| X following graph | 142 | 48.6% |
Our Common Crawl study found that tier's accounts fill in a bio 13 fewer points than curated accounts (81.4% vs. 91.8%) and carry X's blue check a quarter as often. On every signal that study measured, Common Crawl looked less complete, less notable, more ordinary. Location breaks the streak: Common Crawl accounts (64.8%) list it more often than Wikidata's curated, notable ones (51.2%) — a 13.6-point gap in the opposite direction from bio.
The likely reason isn't that Common Crawl accounts are more diligent. It's what each tier actually is. Wikidata seeds this index with people notable enough for an encyclopedia entry — politicians, executives, athletes — many of whom run their account through a comms team or an assistant, or simply never touched the profile-setup fields beyond a name and a bio written for public consumption. A location invites doxxing risk a public figure has no reason to accept. Common Crawl, by contrast, discovers accounts because some ordinary web page — a personal site, a forum profile, a "follow us" footer — already links to them; the same everyday habit of putting yourself on the map that gets you discovered by a link-following crawl also inclines you to type a city into a profile field. GitHub users, unsurprisingly, top every tier (76.8%) — a location line is a normal part of a professional developer profile, the same instinct that makes a résumé list a city.
The familiar gradient still holds — just weaker
Reach still predicts profile completeness, the way it did for bio and verification in earlier studies — the tier reversal above is about what kind of account fills in location, not whether more followers correlates with filling in more fields at all:
| Follower band | Sample size | Location fill rate |
|---|---|---|
| Under 100 | 706 | 55.0% |
| 100–999 | 420 | 67.9% |
| 1,000–9,999 | 419 | 69.2% |
| 10,000–99,999 | 413 | 73.6% |
| 100,000+ | 391 | 69.3% |
| Verification status | Sample size | Location fill rate |
|---|---|---|
| Verified (blue check) | 531 | 73.1% |
| Not verified | 1,931 | 60.7% |
Fill rate climbs 18.6 points from the smallest accounts to the low-tens-of-thousands range, then flattens — the same "diminishing but real" shape the follow-limit study and the bio study found for other profile fields. Verified accounts sit 12.4 points above unverified ones, in line with verification's usual role in this series as a rough proxy for how seriously an account treats its public profile.
Where accounts say they are
Restricting to the 849 locations (54.4% of all filled-in entries) we could confidently match to a single country by keyword — the rest are city-only mentions without a clear country marker, regional/continental answers ("Europe", "Africa"), or genuinely not a place at all (see below):
| Country | Share of classified locations |
|---|---|
| United States | 23.2% |
| Japan | 6.8% |
| United Kingdom | 5.7% |
| Brazil | 2.4% |
| India | 2.4% |
| Germany | 2.2% |
| France | 2.0% |
| Canada | 1.7% |
| Spain | 1.1% |
| Australia | 1.1% |
The most common single strings, unclassified into countries, were exactly the cities you'd expect an English-language, tech- and media-skewed corpus to produce: San Francisco (17 mentions), Los Angeles (14), Washington DC (13), London (12), New York (12) — alongside 13 profiles that wrote 日本 (Japan) in kanji rather than English. The U.S. skew is expected for a corpus seeded partly from an English-language crawl and GitHub, whose global developer base still concentrates in the U.S.; Japan's second-place showing was the bigger surprise, and tracks with a broader pattern in the underlying dataset — this same index's TikTok-creator and Common Crawl tiers also over-index on Japanese and Japanese-language accounts relative to their overall size.
The location field as bio, business card, and inside joke
Not every filled-in location is a place. Across the 1,561 non-empty entries:
| Pattern | Share | Count |
|---|---|---|
| Contains an emoji | 4.3% | 67 |
| A joke, non-place answer ("Earth", "Metaverse", "Everywhere", "Internet") | 1.7% | 27 |
| Raw GPS coordinates instead of a place name | 0.4% | 6 |
| Contact info or business-inquiry text | 1.0% | 15 |
The best example of all belongs to a company, not a person: @Meta, 9.9 million followers, lists its X location as "Metaverse." It isn't alone — the developer accounts @jradoff (303,743 followers) and @thealphadex (38,032 followers) both did the same, unprompted, without any obvious coordination. "Earth" shows up 6 separate times across our sample, including on @AFPphoto, a 202,037-follower news wire account that instead opted for "worldwide." A handful of accounts turned the field into a business card outright — a Japanese illustration account listing a Gmail address and "DM for commission inquiries" where a city should go — and six wrote raw latitude/longitude pairs, one prefixed with the German abbreviation "ÜT" (Übersetzung/translated location — likely an artifact of a non-English X client), including one that pinpoints a spot in Amsterdam down to six decimal places.
What this shows, and what it doesn't
What holds up: location is the one profile-completeness signal in this series where the "ordinary" Common Crawl tier beats the "notable" Wikidata tier, and GitHub's professional-network norms produce the highest fill rate of any tier measured. The reach gradient (more followers, more verification → more complete profile) still holds for location, just less steeply than it did for bio or verification.
What limits it:
- This is a sample, not a census. Unlike this series' other studies, location isn't exposed by the bulk search API — it only appears in a per-account live profile lookup, so we drew a stratified sample of 2,482 accounts (across 8 tiers and 6 follower bands) rather than querying the full 820,548-profile index. At roughly 590–600 accounts per major tier, expect a margin of a few percentage points on any single tier's fill rate, not census-level precision.
- Country classification is loose, keyword-based text matching, not geocoding. We could confidently assign a country to 54.4% of filled-in locations; the rest are ambiguous city names without a country marker, regional answers, non-English scripts our keyword list didn't cover, or genuinely not a place. Treat the country table as a floor on real geographic concentration, not a complete map.
- Location is unverified, user-typed free text. X does nothing to confirm it's real, current, or even a place — which is exactly what makes the joke-answer and coordinate findings possible, but also means every number here describes what accounts say, not where they verifiably are.
- The corpus itself is seeded and notability-skewed, the same caveat as every study in this series — see the Common Crawl piece for what that does and doesn't change about the underlying population.
Sources
Methodology
We drew a stratified sample of usernames from the X users dataset's live search API, split across 8 source tiers (targeting proportionally more from the two largest, Wikidata and Common Crawl) and 6 log-spaced follower bands per tier (0–9, 10–99, 100–999, 1,000–9,999, 10,000–99,999, 100,000+), to avoid oversampling either mega-accounts or the long tail. That produced 2,482 unique usernames. We then queried the live x_profile endpoint individually for each — the only place this dataset's location field is exposed — recording location, current follower count, and verification status for each successful lookup (2,462 of 2,482, a 99.2% resolve rate; the rest were private, suspended, or timed out).
Country classification used a hand-built keyword list (country names, demonyms, and roughly 25 major cities per country, plus a U.S.-state-abbreviation regex for City, ST-style entries) matched case-insensitively against each location string — deliberately conservative, so entries we couldn't confidently match were left unclassified rather than guessed. Junk-pattern detection used a Unicode symbol-category check for emoji, a decimal-coordinate regex for GPS strings, and an exact-match list for common joke answers.
Want to run your own cut? The X users dataset is queryable directly, the how to scrape Twitter/X guide covers both the bulk search and the per-profile x_profile calls used here, and pricing has the free tier. See also the rest of this series: the most prolific poster on X is a machine, X's real follow limit is 7,500, a random developer is 5x more likely than a notable person to have a blue check, what X bios reveal, how the Common Crawl tier changed our X corpus, and X's attention economy is more unequal than any country's wealth.
Frequently asked questions
What share of X profiles fill in a location?
63.4% in our sample — 1,561 of 2,462 successfully-looked-up profiles had a non-empty location field. Fill rate climbs with reach: 55.0% under 100 followers, up to 69–74% above 1,000 followers, and 73.1% for verified accounts versus 60.7% for unverified ones.
Do notable X accounts fill in their location more than ordinary ones?
No — and that's the surprise. Wikidata's curated, notable-people tier fills in location 51.2% of the time, the lowest of any major tier we sampled. Our Common Crawl tier, discovered through ordinary open-web mentions rather than notability, fills it in 64.8% of the time — a 13.6-point gap in the opposite direction from bio and verification, both of which run higher for the curated tier.
Why isn't location in Crawlora's bulk X users search API?
It simply isn't stored in the bulk search/item records (which carry username, bio, follower counts, verification and source tier) — only a live, per-account profile lookup (the x_profile endpoint) returns it. That's why this study used a 2,482-account stratified sample with individual lookups rather than a full-corpus query like the rest of this series.
What weird answers do X accounts put in their location field?
4.3% of filled-in locations contain an emoji, 1.7% are a joke non-place answer like "Earth," "Metaverse," "Everywhere" or "Internet," and 0.4% are raw GPS coordinates instead of a place name. The company account @Meta (9.9 million followers) lists its location as "Metaverse."
Which countries do the most X accounts in this corpus say they're from?
The United States leads at 23.2% of the locations we could confidently classify to a country, followed by Japan (6.8%), the United Kingdom (5.7%), Brazil, India, Germany, France and Canada. We could only classify 54.4% of filled-in locations with confidence — the rest are ambiguous city names, regional answers, or non-place text.
Is this location data representative of all X users?
No, for two compounding reasons. First, the underlying corpus is seeded from notable/curated sources plus open-web mentions, not a random sample of X's 600M+ users. Second, this specific study draws a 2,482-account stratified sample (not a full census) because location requires an individual profile lookup per account rather than a bulk query, so tier-level shares carry a sampling margin of a few points rather than census-level precision.