Social intelligence · X (Twitter) · July 27, 2026
796,611 public X profiles — followers, bio, external link, blue check and account age — queryable over one REST API. Filter by follower band, verification, ratio or creation date; sort by reach. Read logged-out from public pages, no login and no API key, pay on success.
796,611
X profiles — one record per username, fully refreshed this week.
53.7%
1K+ followers
13.2%
blue check
87.6%
have a bio
Snapshot July 27, 2026 — public fields only. A seeded index of established accounts, not a sample of X.
796,611
public X profiles, one record per username, read from the logged-out page — no login and no API key. 46.4% were seeded from Wikidata, so this is an index of established accounts, not a sample of X.
53.7%
have 1,000+ followers — the opposite of a typical social long tail. The seeding skews the population upward, which is the point if you want reachable accounts, and a caveat if you want a census.
13.2%
carry a blue check (105,212 accounts), but it scales hard with reach: 1.9% of sub-100-follower accounts versus 57.3% of the 100K+ tier.
31.6%
were created before 2010, peaking at 164,905 in 2009 alone. The index reaches back to jack and the other founder accounts of March 2006.
Every number on this page describes the accounts in the index, not X's user base. Accounts arrive through seed tiers, and the two largest are now close in size: Wikidata (notable people with an entry and a linked handle) and Common Crawl (accounts surfaced from web-archive mentions, a broader and less curated net). The Wikidata, GitHub and founder tiers skew the population established, old and verified; Common Crawl pulls in a wider slice. It is the right shape for finding and enriching accounts that matter; it is the wrong shape for claims like 'the average X user'.
A random sample of X would be overwhelmingly tiny accounts. This one is not: the modal band is 1K–9.9K followers, 53.7% clear 1,000 followers, and 51,761 accounts sit above 100K — up to 241M on elonmusk. 8,236 accounts have a million or more. The floor is real too: 19,130 records have zero followers, and 254,418 follow more accounts than follow them back — max_ratio is the lever for that shape.
| Followers | Accounts | Share | Verified |
|---|---|---|---|
| Under 100 | 156,174 | 19.6% | 1.9% |
| 100–999 | 212,713 | 26.7% | 4.7% |
| 1K–9.9K | 243,197 | 30.5% | 10.8% |
| 10K–99K | 132,766 | 16.7% | 27.3% |
| 100K+ | 51,761 | 6.5% | 57.3% |
The check has been purchasable since late 2022, so it no longer certifies who someone is. What it does track is reach: the share of accounts carrying it climbs steadily with follower count, from a rounding error at the bottom to a majority of the 100K+ tier. Treat it as a paid-subscription signal that correlates with audience size.
Creation dates cluster hard in the Twitter land-rush: 164,905 of the indexed accounts (20.7%) were created in 2009 alone, and 31.6% predate 2010. That is the Wikidata seeding showing through — notable people signed up early and stayed. The tail thins every year after, though it never disappears.
A profile is only useful if the fields you need are populated. Bios are nearly universal, and three quarters of accounts publish an outbound link — which is what makes this workable as a discovery surface for sites, newsletters and shops behind a handle.
Straight from a followers_desc query — nothing hand-picked. Every one of them registered between 2007 and 2010.
| Username | Name | Followers | Joined |
|---|---|---|---|
@elonmusk | Elon Musk | 240,878,023 | 2009 |
@BarackObama | Barack Obama | 119,205,038 | 2007 |
@realDonaldTrump | Donald J. Trump | 111,753,011 | 2009 |
@Cristiano | Cristiano Ronaldo | 111,749,227 | 2010 |
@narendramodi | Narendra Modi | 107,046,485 | 2009 |
@rihanna | Rihanna | 97,858,371 | 2009 |
@NASA | NASA | 92,203,313 | 2007 |
@justinbieber | Justin Bieber | 91,082,222 | 2009 |
@katyperry | KATY PERRY | 88,104,191 | 2009 |
@taylorswift13 | Taylor Swift | 81,959,294 | 2008 |
@ladygaga | Lady Gaga | 73,490,829 | 2008 |
@imVkohli | Virat Kohli | 70,600,064 | 2009 |
One record per username. Grouped for readability; the API returns a flat object. No post bodies — for those, use the live X endpoints.
Identity
usernameidnameavatar_urlbanner_urlProfile
biolocation_rawexternal_urlhas_biohas_external_urlAudience
followersfollowingfollower_following_ratiopostsStatus & provenance
is_blue_verifiedcreated_atsource_tiercrawled_atschema_versionThe search endpoint takes the filters below; combine any of them. Page with page and page_size (≤100 per page, and page × page_size ≤ 10,000).
Full-text & identity
Audience
Profile signals
Dates
Sort
relevancefollowers_descfollowers_asccreated_at_desccreated_at_asccrawled_at_desccrawled_at_ascA recurring question: how many people on X mention working at an AI lab? The honest answer is a lesson in the search endpoint's limits as much as a number. q is OR-token, not phrase search, so a quoted two-word name like "Meta AI" matches any bio containing Meta OR AI and returns noise. A single distinctive token works — here is every clean token we tested, read by hand for how many hits are genuinely lab-related.
| Token (q=) | Bios matching | Blue-check rate | Reading the hits |
|---|---|---|---|
OpenAI | 373 | 68.6% | Nearly all hits are the official account or a role at the company — the cleanest token tested. |
DeepMind | 150 | 49.3% | Nearly all hits are research staff naming Google DeepMind; lower verification reflects academics, not fewer real employees. |
xAI | 79 | 69.6% | Mostly genuine, with real collisions — a Japanese fan handle and one academic use of "XAI" (explainable AI). |
Anthropic | 51 | 68.6% | About three-quarters staff, ex-staff or investors; the rest are journalists, joke accounts, and one ecology researcher's unrelated use of the word. |
Cohere | 31 | 67.7% | Mostly genuine, minus two unrelated companies sharing the name — a wireless-tech firm and a coworking brand. |
Mistral | 24 | 62.5% | Under three-quarters genuine; the rest are a French bus network, a university and bookstore named for poet Gabriela Mistral, and a restaurant. |
DeepSeek | 10 | 80% | Clean token, but only ~2 of 10 hits are the company or an employee — the rest mention the product as a tool. |
Qwen | 8 | 50% | Clean token, but mostly users running the model or program "ambassadors," not employees. |
These are bio-mention counts, not an employee census — most bios never name an employer at all, and Kimi / Moonshot are excluded here because reading their hits shows the tokens are dominated by unrelated collisions, not the AI labs.
Read the full study, methodology and caveats →Query a single distinctive token
# Bio mentions of a single distinctive AI-lab name — quoting a two-word
# name like "Meta AI" does NOT work here (q is OR-token, not phrase search)
curl "https://api.crawlora.net/api/v1/datasets/x-users/search?q=DeepMind&has_bio=true&page_size=30" \
-H "x-api-key: $CRAWLORA_API_KEY"Every query authenticates with an x-api-key header and reads the stored search index — there is no live crawl or proxy to manage, and you are billed pay on success: charged for results, not failed requests. Three endpoints cover it:
GET /datasets/x-users/search — filter, sort, page the population.GET /datasets/x-users/items/{username} — one profile.GET /datasets/x-users/facets — counts for any dimension.The same three are exposed as MCP tools — datasets_x_users_search, datasets_x_users_item and datasets_x_users_facets — so an agent can call them directly. Need post bodies, a profile timeline or a live read of an account that is not in the index? Those are the live X endpoints.
Cite this
Crawlora (2026). X Users Dataset. 796,611 public X profiles, seeded index; public fields only. https://crawlora.net/datasets/x-users.
Search with filters
# Verified accounts with 100K+ followers, biggest first
curl "https://api.crawlora.net/api/v1/datasets/x-users/search?is_blue_verified=true&min_followers=100000&sort=followers_desc" \
-H "x-api-key: $CRAWLORA_API_KEY"One profile by username
# One profile by username (leading @ optional)
curl "https://api.crawlora.net/api/v1/datasets/x-users/items/elonmusk" -H "x-api-key: $CRAWLORA_API_KEY"Facet the population
# Facet the population by seed tier (or is_blue_verified, has_bio, has_external_url)
curl "https://api.crawlora.net/api/v1/datasets/x-users/facets?facet=source_tier" -H "x-api-key: $CRAWLORA_API_KEY"Filter by follow ratio
# Accounts created since 2023 that follow far more than follow back
curl "https://api.crawlora.net/api/v1/datasets/x-users/search?created_after=2023-01-01&max_ratio=0.5" \
-H "x-api-key: $CRAWLORA_API_KEY"The same figures behind the bars and the line — plain and machine-readable for search engines and AI answer engines that cannot parse a chart.
| Created | Accounts | Share of index |
|---|---|---|
| 2006 | 1,662 | 0.2% |
| 2007 | 26,988 | 3.4% |
| 2008 | 58,107 | 7.3% |
| 2009 | 164,905 | 20.7% |
| 2010 | 85,816 | 10.8% |
| 2011 | 81,472 | 10.2% |
| 2012 | 56,243 | 7.1% |
| 2013 | 43,953 | 5.5% |
| 2014 | 38,347 | 4.8% |
| 2015 | 34,559 | 4.3% |
| 2016 | 30,787 | 3.9% |
| 2017 | 29,185 | 3.7% |
| 2018 | 26,956 | 3.4% |
| 2019 | 24,038 | 3% |
| 2020 | 19,712 | 2.5% |
| 2021 | 17,526 | 2.2% |
| 2022 | 16,520 | 2.1% |
| 2023 | 11,797 | 1.5% |
| 2024 | 10,806 | 1.4% |
| 2025 | 11,047 | 1.4% |
| 2026 (partial) | 6,185 | 0.8% |
| Followers | Accounts | Share | Verified | Verified % |
|---|---|---|---|---|
| Under 100 | 156,174 | 19.6% | 2,970 | 1.9% |
| 100–999 | 212,713 | 26.7% | 9,962 | 4.7% |
| 1K–9.9K | 243,197 | 30.5% | 26,366 | 10.8% |
| 10K–99K | 132,766 | 16.7% | 36,238 | 27.3% |
| 100K+ | 51,761 | 6.5% | 29,676 | 57.3% |
| Seed tier | Accounts | Share |
|---|---|---|
| Wikidata (notable people) | 369,977 | 46.4% |
| Common Crawl | 324,523 | 40.7% |
| GitHub profiles | 42,797 | 5.4% |
| X search | 22,367 | 2.8% |
| X following graph | 14,696 | 1.8% |
| Unknown | 12,070 | 1.5% |
| TrustMRR startups | 4,334 | 0.5% |
| Brave search | 1,725 | 0.2% |
| TikTok creators | 1,476 | 0.2% |
| Journalists | 1,257 | 0.2% |
| Manual seeds | 232 | 0% |
| X follower graph | 59 | 0% |
| Chrome search | 57 | 0% |
| X page | 2 | 0% |
Query 797K public X profiles — followers, bio, external link, verification and account age — through one REST API. Creator discovery, audience research, lead enrichment or account vetting: clean JSON, public fields only, pay on success.
The dataset holds 796,611 public X (Twitter) profiles, deduplicated to one record per username. Every boolean facet sums to that exact total. 1,039 records predate source-tier tracking and carry no tier field at all, so they drop out of the source_tier breakdown rather than landing in the "unknown" bucket — which is why that one column sums slightly short.
No, and it matters. This is a seeded index: accounts are discovered through seed tiers — 46.4% from Wikidata (notable people), 5.4% from GitHub profiles that link an X handle, then X search, the following graph, TikTok creators, journalists and others. That seeding pulls the population toward established accounts: 53.7% have 1,000+ followers and 31.6% were created before 2010, neither of which is true of X as a whole. Use it as a directory of notable and reachable accounts, not as a census or a basis for population-level claims about X users.
Every field is a public, credential-free profile field read from the logged-out profile page: username, id, display name, avatar and banner URLs, bio, self-declared location, external link, follower and following counts, the follower-to-following ratio, post count, the blue-check flag, account creation date, the seed tier that discovered the account, and a crawled_at timestamp. Nothing behind authentication is collected — no login, no API key, no guest token — and protected accounts are never indexed. There are no email addresses, no direct messages and no post bodies; for posts, use the live X endpoints instead.
105,212 accounts (13.2%) carry the blue check, but it tracks reach closely: 1.9% of accounts under 100 followers are verified, rising through 10.8% in the 1K–9.9K band to 57.3% of the 100K+ tier. Since the check has been purchasable since 2022 it reads as a paid-subscription signal that correlates with reach, not as an identity guarantee — filter on it accordingly.
Very. Every one of the 796,611 records in this snapshot was re-crawled between July 14, 2026 and July 27, 2026 — a full refresh of the index in an eight-day window, not a backfill with a stale tail. Each record carries its own crawled_at timestamp, and you can filter or sort on it (crawled_after, crawled_before, sort=crawled_at_asc) to find the stalest records yourself.
Yes. Three endpoints cover it: search (full-text and faceted, with follower, ratio, date and profile-signal filters), one profile by username, and facet counts for any dimension. The same three are exposed as MCP tools — datasets_x_users_search, datasets_x_users_item and datasets_x_users_facets — so an agent can call them directly. Reads hit the stored index and never trigger a live crawl; billing is pay-on-success.