Social intelligence · X (Twitter) · July 21, 2026
471,437 public X profiles — followers, bio, external link, blue check and account age — queryable over one REST API. Filter by follower band, verification, ratio or creation date; sort by reach. Read logged-out from public pages, no login and no API key, pay on success.
471,437
X profiles — one record per username, fully refreshed this week.
66.1%
1K+ followers
18.9%
blue check
91.8%
have a bio
Snapshot July 21, 2026 — public fields only. A seeded index of established accounts, not a sample of X.
471,437
public X profiles, one record per username, read from the logged-out page — no login and no API key. 78.5% were seeded from Wikidata, so this is an index of established accounts, not a sample of X.
66.1%
have 1,000+ followers — the opposite of a typical social long tail. The seeding skews the population upward, which is the point if you want reachable accounts, and a caveat if you want a census.
18.9%
carry a blue check (89,172 accounts), but it scales hard with reach: 4.5% of sub-100-follower accounts versus 57.8% of the 100K+ tier.
21.6%
were created before 2010, peaking at 79,024 in 2009 alone. The index reaches back to jack and the other founder accounts of March 2006.
Every number on this page describes the accounts in the index, not X's user base. Accounts arrive through seed tiers, and the biggest tier by far is Wikidata — notable people with an entry and a linked handle. That is why the population skews established, old and verified. It is the right shape for finding and enriching accounts that matter; it is the wrong shape for claims like 'the average X user'.
A random sample of X would be overwhelmingly tiny accounts. This one is not: the modal band is 1K–9.9K followers, 66.1% clear 1,000 followers, and 46,410 accounts sit above 100K — up to 241M on elonmusk. 7,877 accounts have a million or more. The floor is real too: 5,440 records have zero followers, and 97,576 follow more accounts than follow them back — max_ratio is the lever for that shape.
| Followers | Accounts | Share | Verified |
|---|---|---|---|
| Under 100 | 59,410 | 12.6% | 4.5% |
| 100–999 | 100,490 | 21.3% | 8.2% |
| 1K–9.9K | 158,244 | 33.6% | 13.2% |
| 10K–99K | 106,883 | 22.7% | 28.6% |
| 100K+ | 46,410 | 9.8% | 57.8% |
The check has been purchasable since late 2022, so it no longer certifies who someone is. What it does track is reach: the share of accounts carrying it climbs steadily with follower count, from a rounding error at the bottom to a majority of the 100K+ tier. Treat it as a paid-subscription signal that correlates with audience size.
Creation dates cluster hard in the Twitter land-rush: 79,024 of the indexed accounts (16.8%) were created in 2009 alone, and 21.6% predate 2010. That is the Wikidata seeding showing through — notable people signed up early and stayed. The tail thins every year after, though it never disappears.
A profile is only useful if the fields you need are populated. Bios are nearly universal, and three quarters of accounts publish an outbound link — which is what makes this workable as a discovery surface for sites, newsletters and shops behind a handle.
Straight from a followers_desc query — nothing hand-picked. Every one of them registered between 2007 and 2010.
| Username | Name | Followers | Joined |
|---|---|---|---|
@elonmusk | Elon Musk | 240,878,023 | 2009 |
@BarackObama | Barack Obama | 119,205,038 | 2007 |
@realDonaldTrump | Donald J. Trump | 111,753,011 | 2009 |
@Cristiano | Cristiano Ronaldo | 111,749,227 | 2010 |
@narendramodi | Narendra Modi | 107,046,485 | 2009 |
@rihanna | Rihanna | 97,858,371 | 2009 |
@NASA | NASA | 92,203,313 | 2007 |
@justinbieber | Justin Bieber | 91,082,222 | 2009 |
@katyperry | KATY PERRY | 88,104,191 | 2009 |
@taylorswift13 | Taylor Swift | 81,959,294 | 2008 |
@ladygaga | Lady Gaga | 73,490,829 | 2008 |
@imVkohli | Virat Kohli | 70,600,064 | 2009 |
One record per username. Grouped for readability; the API returns a flat object. No post bodies — for those, use the live X endpoints.
Identity
usernameidnameavatar_urlbanner_urlProfile
biolocation_rawexternal_urlhas_biohas_external_urlAudience
followersfollowingfollower_following_ratiopostsStatus & provenance
is_blue_verifiedcreated_atsource_tiercrawled_atschema_versionThe search endpoint takes the filters below; combine any of them. Page with page and page_size (≤100 per page, and page × page_size ≤ 10,000).
Full-text & identity
Audience
Profile signals
Dates
Sort
relevancefollowers_descfollowers_asccreated_at_desccreated_at_asccrawled_at_desccrawled_at_ascEvery query authenticates with an x-api-key header and reads the stored search index — there is no live crawl or proxy to manage, and you are billed pay on success: charged for results, not failed requests. Three endpoints cover it:
GET /datasets/x-users/search — filter, sort, page the population.GET /datasets/x-users/items/{username} — one profile.GET /datasets/x-users/facets — counts for any dimension.The same three are exposed as MCP tools — datasets_x_users_search, datasets_x_users_item and datasets_x_users_facets — so an agent can call them directly. Need post bodies, a profile timeline or a live read of an account that is not in the index? Those are the live X endpoints.
Cite this
Crawlora (2026). X Users Dataset. 471,437 public X profiles, seeded index; public fields only. https://crawlora.net/datasets/x-users.
Search with filters
# Verified accounts with 100K+ followers, biggest first
curl "https://api.crawlora.net/api/v1/datasets/x-users/search?is_blue_verified=true&min_followers=100000&sort=followers_desc" \
-H "x-api-key: $CRAWLORA_API_KEY"One profile by username
# One profile by username (leading @ optional)
curl "https://api.crawlora.net/api/v1/datasets/x-users/items/elonmusk" -H "x-api-key: $CRAWLORA_API_KEY"Facet the population
# Facet the population by seed tier (or is_blue_verified, has_bio, has_external_url)
curl "https://api.crawlora.net/api/v1/datasets/x-users/facets?facet=source_tier" -H "x-api-key: $CRAWLORA_API_KEY"Filter by follow ratio
# Accounts created since 2023 that follow far more than follow back
curl "https://api.crawlora.net/api/v1/datasets/x-users/search?created_after=2023-01-01&max_ratio=0.5" \
-H "x-api-key: $CRAWLORA_API_KEY"The same figures behind the bars and the line — plain and machine-readable for search engines and AI answer engines that cannot parse a chart.
| Created | Accounts | Share of index |
|---|---|---|
| 2006 | 359 | 0.1% |
| 2007 | 5,178 | 1.1% |
| 2008 | 17,068 | 3.6% |
| 2009 | 79,024 | 16.8% |
| 2010 | 54,085 | 11.5% |
| 2011 | 51,656 | 11% |
| 2012 | 39,726 | 8.4% |
| 2013 | 31,002 | 6.6% |
| 2014 | 26,340 | 5.6% |
| 2015 | 23,102 | 4.9% |
| 2016 | 19,230 | 4.1% |
| 2017 | 18,102 | 3.8% |
| 2018 | 16,571 | 3.5% |
| 2019 | 15,581 | 3.3% |
| 2020 | 16,440 | 3.5% |
| 2021 | 14,518 | 3.1% |
| 2022 | 13,100 | 2.8% |
| 2023 | 9,202 | 2% |
| 2024 | 8,265 | 1.8% |
| 2025 | 8,082 | 1.7% |
| 2026 (partial) | 4,806 | 1% |
| Followers | Accounts | Share | Verified | Verified % |
|---|---|---|---|---|
| Under 100 | 59,410 | 12.6% | 2,690 | 4.5% |
| 100–999 | 100,490 | 21.3% | 8,204 | 8.2% |
| 1K–9.9K | 158,244 | 33.6% | 20,877 | 13.2% |
| 10K–99K | 106,883 | 22.7% | 30,583 | 28.6% |
| 100K+ | 46,410 | 9.8% | 26,818 | 57.8% |
| Seed tier | Accounts | Share |
|---|---|---|
| Wikidata (notable people) | 369,977 | 78.5% |
| GitHub profiles | 42,797 | 9.1% |
| X search | 22,367 | 4.7% |
| X following graph | 14,048 | 3% |
| Unknown | 12,070 | 2.6% |
| TrustMRR startups | 4,334 | 0.9% |
| Brave search | 1,725 | 0.4% |
| TikTok creators | 1,476 | 0.3% |
| Journalists | 1,257 | 0.3% |
| Manual seeds | 232 | 0% |
| X follower graph | 59 | 0% |
| Chrome search | 57 | 0% |
| X page | 2 | 0% |
Query 471K public X profiles — followers, bio, external link, verification and account age — through one REST API. Creator discovery, audience research, lead enrichment or account vetting: clean JSON, public fields only, pay on success.
The dataset holds 471,437 public X (Twitter) profiles, deduplicated to one record per username. Every boolean facet sums to that exact total. 1,036 records predate source-tier tracking and carry no tier field at all, so they drop out of the source_tier breakdown rather than landing in the "unknown" bucket — which is why that one column sums slightly short.
No, and it matters. This is a seeded index: accounts are discovered through seed tiers — 78.5% from Wikidata (notable people), 9.1% from GitHub profiles that link an X handle, then X search, the following graph, TikTok creators, journalists and others. That seeding pulls the population toward established accounts: 66.1% have 1,000+ followers and 21.6% were created before 2010, neither of which is true of X as a whole. Use it as a directory of notable and reachable accounts, not as a census or a basis for population-level claims about X users.
Every field is a public, credential-free profile field read from the logged-out profile page: username, id, display name, avatar and banner URLs, bio, self-declared location, external link, follower and following counts, the follower-to-following ratio, post count, the blue-check flag, account creation date, the seed tier that discovered the account, and a crawled_at timestamp. Nothing behind authentication is collected — no login, no API key, no guest token — and protected accounts are never indexed. There are no email addresses, no direct messages and no post bodies; for posts, use the live X endpoints instead.
89,172 accounts (18.9%) carry the blue check, but it tracks reach closely: 4.5% of accounts under 100 followers are verified, rising through 13.2% in the 1K–9.9K band to 57.8% of the 100K+ tier. Since the check has been purchasable since 2022 it reads as a paid-subscription signal that correlates with reach, not as an identity guarantee — filter on it accordingly.
Very. Every one of the 471,437 records in this snapshot was re-crawled between July 14, 2026 and July 21, 2026 — a full refresh of the index in an eight-day window, not a backfill with a stale tail. Each record carries its own crawled_at timestamp, and you can filter or sort on it (crawled_after, crawled_before, sort=crawled_at_asc) to find the stalest records yourself.
Yes. Three endpoints cover it: search (full-text and faceted, with follower, ratio, date and profile-signal filters), one profile by username, and facet counts for any dimension. The same three are exposed as MCP tools — datasets_x_users_search, datasets_x_users_item and datasets_x_users_facets — so an agent can call them directly. Reads hit the stored index and never trigger a live crawl; billing is pay-on-success.