Übersetzung in Arbeit. Diese Seite wird vorübergehend auf Englisch angezeigt, während die deutsche Übersetzung vorbereitet wird.
Data layer for AI agents
Ready-made web datasets, queryable by API.
Structured, ready-made web data your AI agents and pipelines query over one REST API — or straight from an MCP client: a 5.2M-app App Store + Google Play catalog, the daily app charts for both stores, 15M app reviews with sentiment, a Google Maps local-business dataset, and live web studies like the Anti-Bot and Dead-Web indexes. Pull clean JSON, pay on success — no crawl to run.
5.2M
apps indexed, plus 42 more datasets — all one REST API away.
Every app's category, rating, ratings count, price, install scale and popularity
5.2M apps
Continuously enriched
datasets/apps/search
App Store + Google Play charts
The daily Top Free, Top Paid and Top Grossing rankings on the iOS App Store and Google Play
89 charts daily
Refreshed daily
datasets/apps-charts/search
App Store + Google Play reviews
Recent user reviews of the top apps on the iOS App Store and Google Play
15M reviews
Periodically refreshed
datasets/apps-reviews/search
Google Maps businesses
Scraped Google Maps business records held in a search index
Local business listings
Continuously expanded
datasets/google-map-businesses/search
Airbnb markets
Aggregate Airbnb short-term-rental market statistics from a near-complete public-listing supply census
8.5M listings
Snapshot 2026-07-03
datasets/airbnb-markets/search
GitHub developer profiles
Enriched public profiles for individual GitHub developers
1.6M profiles
Snapshot 2026-07-07
datasets/github-users/search
X (Twitter) user profiles
Public X (Twitter) profiles
821K profiles
Snapshot 2026-08-30
datasets/x-users/search
Instagram user profiles
Public Instagram profiles
617K profiles
Snapshot 2026-08-29
datasets/instagram-users/search
YouTube channel profiles
Public YouTube channel profiles
4.6M channels
Snapshot 2026-08-29
datasets/youtube-creators/search
Facebook Page contact records
Public Facebook Page contact records
1.3M Pages
Snapshot 2026-08-30
datasets/facebook-pages/search
BBB business profiles
Better Business Bureau business profiles
119K profiles
Snapshot 2026-08-30
datasets/bbb-businesses/search
Verified-revenue startups
Public startups with payment-provider-verified revenue and MRR
10K startups
Snapshot 2026-08-30
datasets/trustmrr/search
Box office charts
Theatrical box-office records from public Box Office Mojo charts and title pages
22K titles
Snapshot 2026-08-29
datasets/boxofficemojo/search
PlayStation Store catalog
Crawled public PlayStation Store catalog
134K SKUs
Snapshot 2026-08-29
datasets/playstation-games/search
Used-vehicle listings
Crawled used-vehicle listings across CarMax, Autotrader and Cars.com
196K listings
Snapshot 2026-08-30
datasets/vehicle-listings/search
Apple Podcasts shows
Crawled public Apple Podcasts show catalog
890K shows
Snapshot 2026-08-29
datasets/apple-podcasts-shows/search
Goodreads books & authors
Crawled public Goodreads book catalog
22M books
Snapshot 2026-08-29
datasets/goodreads-books/search
Product Hunt launches & makers
Individual Product Hunt launches with topics, upvotes, ranks and launch history, maker profiles for leaderboards, and aggregate topic/period trend rollups
120K launches
Snapshot 2026-08-30
datasets/producthunt-products/search
PitchBook private markets
Crawled public PitchBook company, fund, investor, advisor and limited-partner profile catalogs
1.4M companies
Snapshot 2026-08-30
datasets/pitchbook-companies/search
SEC EDGAR companies
Every SEC EDGAR filer as one record
977K filers
Snapshot 2026-08-30
datasets/sec-companies/search
Jobs — companies & postings
Discovered company career-site boards across 16 ATS providers, plus 5 single-company big-tech careers platforms (Amazon, Apple, Google, Meta, Tesla)
70K boards
Snapshot 2026-09-05
datasets/jobs/search
Steam games
The top 10,000 Steam games by SteamSpy owner estimate
10K top games
Snapshot 2026-07-11
datasets/steam-games/search
Chrome Web Store items
Extensions, themes and legacy apps with displayed users, ratings, versions, declared permissions, privacy disclosures and change history
966 pilot records
Pilot snapshot 2026-07-14
datasets/chrome-extensions/search
Steam charts
Steam's official store charts as a daily time series
220 rows daily
Refreshed daily
datasets/steam-charts/search
Reddit trending posts
Each tracked subreddit's current hot-feed post order, captured daily and growing toward 10,000 subreddits
100 subreddits tracked
Refreshed daily
datasets/reddit-trending/search
US housing markets
Redfin market statistics for every US metro, county, city and ZIP
52K regions
Snapshot 2026-05
datasets/housing-markets/search
Website tech stack
The technologies behind every site
1.0M sites
Snapshot 2026-08-29
datasets/techstack/search
Tech Stack Index
What the web actually runs on
1,000,000 domains fingerprinted
Snapshot 2026-08-20
—
Journalist contacts
Journalist and reporter records
46K+ journalists
Continuously expanded
datasets/journalists/search
Numbeo cost of living
Cost of living, quality of life, crime & safety, health care, pollution, traffic and property investment
1,083 cities
Snapshot August 29, 2026
datasets/numbeo-cities/search
Most-Played Steam Games
Steam's official most-played chart
Top 100 live
Live
—
Local Business on the Web
How online local businesses really are: 40% of 12.4M US/UK/Canada Google Maps listings have no website, the most common listed “website” is a Facebook page, and coverage runs from 93% (banks) to 33% (farms)
12.4M listings analyzed
Snapshot 2026-06-10
—
SERP Volatility Index
How much Google, Bing and Brave disagree
60 keywords × 3 engines
Daily snapshot
—
Reddit SERP Index
How much of page one Reddit actually owns
260 queries × 4 engines
Snapshot
—
TikTok Creators Index
The TikTok creator pyramid: among 3.26M established creators, only 1.6% have 1M+ followers, 4% are verified, and 4% list a business contact
3.26M creators analyzed
Snapshot 2026-06-21
—
AI Researchers X Index
A roster-first directory of named researchers/staff across 9 frontier AI labs, each resolved to a self-asserted X handle via ORCID, Wikidata, or their own GitHub profile
93 named researchers resolved
Snapshot 2026-07-27
—
Airbnb Market Index
Short-term-rental market data across 10 major Airbnb markets: estimated occupancy runs 22%–67%, nightly rates $69–$464, and professional multi-listing operators range 4%–24%
10 markets analyzed
Snapshot 2026-06-23
—
Anti-Bot Index
Which of the world's top sites run anti-bot / WAF protection, and which vendor
998K sites scanned
Snapshot June 14, 2026
—
Dead-Web Index
How much of the most popular web is actually dead, blocked, or parked
Top global domains, probed
Probed live
—
AI-Crawler Blocking Index
Which AI crawlers each site blocks in robots.txt
Top 1M sites · robots.txt
Snapshot 2026-06-20
—
State of Web Scraping & Anti-Bot 2026
The flagship annual report
Top 1M–10M · synthesis
June 2026
—
Search vs. Store
Do people download the AI apps they Google? Joins Google Trends search demand, US App Store rank and SimilarWeb web traffic for 40 major AI apps
40 AI apps · 3 signals
Snapshot 2026-07-01
—
Search vs. Store — living index
Does Google search demand predict App Store rank? The daily panel behind the study
44 days · 1,760 observations
Daily
—
The app and Google Maps datasets return JSON from the REST endpoints above; the Anti-Bot and Dead-Web indices are interactive studies. Counts are the latest indexed totals and grow as the crawl runs.
The catalog in numbers
What the data looks like.
A few cuts from the two most data-rich datasets — hover any chart to isolate a series. The full breakdowns live on each dataset's page.
The app catalog skews Android
Google Play is ~1.3× the App Store — 3.0M Android apps to 2.3M iOS.
Google Play (Android)3.0M
2,964,796 apps
App Store (iOS)2.3M
2,251,582 apps
Indexed apps per store; the catalog keeps one deduplicated record per app per store.
Most apps are invisible
60.9% of Android apps have under 1,000 installs; only 174 have crossed a billion.
< 1K1,805,589
1K–10K559,564
10K–100K330,934
100K–1M148,666
1M–10M53,392
10M–100M12,428
100M–1B1,313
1B+174
Android apps grouped by Google Play install count — a deep long tail of low-install apps.The same figures as a table
Install bucket
Apps
< 1K
1,805,589
1K–10K
559,564
10K–100K
330,934
100K–1M
148,666
1M–10M
53,392
10M–100M
12,428
100M–1B
1,313
1B+
174
Most of the web doesn’t fight back
Of 998,497 top sites scanned for the Anti-Bot Index, 437,857 run a managed wall — but only 79,835 (8%) actively challenged a bot.
How a million scanned sites narrow to the few that actively fight bots. From the Anti-Bot Index.Show the flows
Stored, structured, and deduplicated — query the index directly without a live crawl. Filter, facet, sort, and page through normalized records over the REST API.
App intelligence
App Store + Google Play apps
5.2M apps
2.3M iOS · 3.0M Android
Every app's category, rating, ratings count, price, install scale and popularity — deduplicated to one record per app per store, across the iOS App Store and Google Play.
The daily Top Free, Top Paid and Top Grossing rankings on the iOS App Store and Google Play — one record per chart, snapshot and rank, so it's a time series you can diff for movers, not just today's leaderboard.
Recent user reviews of the top apps on the iOS App Store and Google Play — score, title, text, helpful count and date, one record per review. Recent-review sentiment runs harsher than lifetime ratings, which is exactly the signal.
Coverage
App Store + Google Play · top apps
Updated
Periodically refreshed · snapshot September 5, 2026
Scraped Google Maps business records held in a search index — query by keyword and location, facet by category, find places near a point, or fetch one business by place_id. Reads the stored index, so there's no live scraping or proxy routing.
Aggregate Airbnb short-term-rental market statistics from a near-complete public-listing supply census — listing supply, ratings and nightly-price bands rolled up by country, metro and geo cell. Aggregate-only: no individual listings or hosts, thin cells suppressed.
Coverage
60+ countries · ~90% of global Airbnb supply
Updated
Snapshot 2026-07-03
Key fields
countrymarketlistingsavg_ratingmedian nightly pricesuperhost sharegeo density
Enriched public profiles for individual GitHub developers — followers, company, interest domains, recent activity and clean reverse-geocoded location, one record per login. Query the stored index directly, no live crawl.
Coverage
Worldwide · individual developers · public fields only
Public X (Twitter) profiles — followers, following, bio, external link, post count, blue check and account age, one record per username. A seeded index of established accounts, read logged-out from public pages; query the stored index directly, no live crawl.
Coverage
Worldwide · seeded index of established accounts, not a sample of X · public fields only
Public Instagram profiles — followers, following, bio, external link, post count, verification and business/creator status, one record per username. A seeded index of established and business accounts, read from public pages; query the stored index directly, no live crawl.
Coverage
Worldwide · seeded index of established and business accounts, not a sample of Instagram · public fields only
Public YouTube channel profiles — subscribers, videos, views, region, bio and links, one record per channel. A long-tail index discovered via Common Crawl and Wikidata, read from each channel's public About page; query the stored index directly, no live crawl or API quota.
Coverage
Worldwide · seeded index discovered via Common Crawl and Wikidata, not a sample of YouTube · public fields only
Public Facebook Page contact records — website, email, phone, WhatsApp, category and like count, one record per Page. A seeded index discovered mostly via business-listing search (Facebook's robots.txt blocks Common Crawl on Page paths), read from each Page's public About tab; query the stored index directly, no live crawl.
Coverage
Worldwide · seeded index discovered mostly via business-listing search, not a sample of Facebook · public fields only
Better Business Bureau business profiles — letter rating, paid-accreditation status, category, legal entity type, location, years in business and public contact details, one record per business. Rating coverage is uneven by design: every accredited business carries a rating, most non-accredited ones do not, so name the denominator before quoting a rating share.
Coverage
United States and Canada · home-services-heavy footprint, not a business registry · public profile fields only
Public startups with payment-provider-verified revenue and MRR — plus traffic, growth, category, tech stack, marketing channels and acquisition-marketplace signals, one record per listing. A marketplace snapshot of the startups TrustMRR itself lists, not a census of every SaaS startup; query the stored index directly, no live crawl.
Coverage
Worldwide · marketplace snapshot of TrustMRR's own listings, not a census · public fields only
Theatrical box-office records from public Box Office Mojo charts and title pages — lifetime and yearly grosses, release groups, market breakdowns, franchise, brand and genre tags. Box Office Mojo's own charted universe, not every film ever released; query the stored index directly, no live crawl.
Coverage
Worldwide · Box Office Mojo's own charted universe, not every film released · public chart data only
Crawled public PlayStation Store catalog — one row per product SKU (game, edition or add-on), with title, publisher, genres, platform, price and star rating. Query the stored index directly, no live crawl or PSN account.
Coverage
US storefront · one row per SKU, not per title · public storefront fields only
Crawled used-vehicle listings across CarMax, Autotrader and Cars.com — price, mileage, VIN, seller type and recorded price-change history, one record per listing. CarMax is a single retailer's full national catalog; Autotrader and Cars.com are best-effort marketplace sweeps, not one uniform census. Query the stored index directly, no live crawl.
Coverage
US · CarMax full catalog plus Autotrader/Cars.com best-effort sweeps, not a uniform census · public fields only
Crawled public Apple Podcasts show catalog — episode count, genre, RSS feed URL, storefront and content rating, one record per show. Discovered from a country x genre x collection chart grid plus a search-term sweep, not a full catalog of every show; query the stored index directly, no live crawl.
Coverage
Chart-grid + search-term discovery, not a full catalog · public fields only
Crawled public Goodreads book catalog — title, authors, series, genres, format, publisher, ISBN and the full ratings/reviews breakdown — plus author profiles discovered as a byproduct of the books crawl. Two datasets, one REST API, no live crawl.
Coverage
Curated Listopia lists + search-term sweep + author bibliography expansion · public fields only
Individual Product Hunt launches with topics, upvotes, ranks and launch history, maker profiles for leaderboards, and aggregate topic/period trend rollups — three datasets, one REST API, pay on success.
Coverage
Worldwide · products, makers and aggregate trend rollups · public fields only
Crawled public PitchBook company, fund, investor, advisor and limited-partner profile catalogs — overview, description, contact/HQ and relationship previews — discovered from PitchBook's own public sitemap. Five datasets, one REST API, no login.
Coverage
Worldwide · PitchBook's own public sitemap, not the paid platform · public fields only
Every SEC EDGAR filer as one record — profile, tickers, industry, filing history, XBRL financial-statement history and insider-transaction rollups — plus institutional 13F holdings by manager. Built from SEC public data, queryable directly, no live crawl.
Coverage
US SEC filers · public EDGAR data · financials where XBRL-reported
Discovered company career-site boards across 16 ATS providers, plus 5 single-company big-tech careers platforms (Amazon, Apple, Google, Meta, Tesla) — live postings from all of them normalized into one JSON shape — company, department, location, employment type, remote flag. Query the stored index directly, no live crawl.
Coverage
Worldwide · 16 ATS providers + 5 big-tech platforms · public postings only
The top 10,000 Steam games by SteamSpy owner estimate — price, review tier, owner band, genres, platforms and release year, one enriched record per app. Query the stored index directly, no live crawl.
Coverage
Worldwide · top 10,000 by owners · public storefront fields
Extensions, themes and legacy apps with displayed users, ratings, versions, declared permissions, privacy disclosures and change history — plus chart-ready aggregate metrics. The current public study is a quality-checked 1,000-ID pilot before the full-catalog backfill.
Steam's official store charts as a daily time series — live concurrent players and most-played by 24h peak (global), plus the US top sellers with price. One record per chart, day and rank, so you can diff for movers, not just read today's leaderboard.
Coverage
Global (players) · US (sales) · Steam official charts
Each tracked subreddit's current hot-feed post order, captured daily and growing toward 10,000 subreddits — one record per subreddit, day and rank, so a subreddit's trending posts can be tracked over time. No score or comment-count field: the credential-free source doesn't expose vote counts.
Redfin market statistics for every US metro, county, city and ZIP — median sale and list price, inventory, days on market and competitiveness, monthly back to 2012 — joined to Census income for price-to-income, salary-to-buy and the affordability gap. Aggregate market data only; reads a stored index, no live scraping.
Coverage
United States · metro/county/city/ZIP · Redfin + Census
The technologies behind every site — CMS, ecommerce platform, CDN, analytics, web server and ~200 others — fingerprinted from HTML + response headers across the Tranco top-1M, one record per domain. Query the stored index directly, no live crawl.
Coverage
Worldwide · Tranco top-1M · public homepage fields only
Journalist and reporter records — outlet, title, beat topics and whatever public contact info (work email or social handle) that outlet lists on its own staff page — crawled directly from news outlets' own public pages, one record per journalist.
Cost of living, quality of life, crime & safety, health care, pollution, traffic and property investment — all seven Numbeo indices merged into one record per city and per country, on the New York = 100 comparable scale. Aggregate index data only; reads a stored index, no live scraping.
Every dataset is a documented REST endpoint. Authenticate with an x-api-key header, filter and sort with query params, and get normalized JSON — no HTML parsing, no proxies to manage. Queries read the stored index, and you’re billed pay on success: charged for results, not failed requests.
Swap in datasets/google-map-businesses/search for local businesses.
Need a sample, or a dataset we don’t list yet?
Crawlora runs 1781+ live endpoints across search, maps, marketplaces, social, finance and more — most public web data can become a structured dataset. Browse what’s live, or tell us what you need.
Crawlora maintains a family of web-data datasets: an App Store + Google Play app catalog (5,216,378 apps — 2,251,582 iOS and 2,964,796 Android), the daily App Store + Google Play top-chart rankings (89 charts across 15 countries), 14,532,024 app reviews with sentiment, a Google Maps local-business dataset, a near-complete Airbnb markets dataset (8.5M listings across 60+ countries), and live web studies such as the Anti-Bot Index (998,497 sites scanned) and the Dead-Web Index. The apps, charts, reviews, Maps and Airbnb datasets are queryable as JSON over the REST API; the indices are interactive studies you can browse and cite.
How do I access the datasets via API?
Query them through Crawlora's REST API with an x-api-key header. `datasets/apps/search` searches the app catalog (filter by store, category, rating, price, sort by popularity) and `datasets/google-map-businesses/search` searches stored Google Maps businesses with keyword, location, facets and nearby queries. Both return normalized JSON and bill pay-on-success. See the Datasets API docs to start.
How fresh are the datasets?
The app catalog is continuously enriched (latest snapshot August 30, 2026) and the Google Maps dataset is continuously expanded. The Anti-Bot Index is a periodic full top-1M scan (latest June 14, 2026); the Dead-Web Index probes domains live on request.
How many apps are in the App Store and Google Play?
Crawlora's catalog has indexed 5,216,378 apps across both stores — 2,251,582 on the iOS App Store and 2,964,796 on Google Play — each deduplicated to one record per store with category, rating, price, install and popularity fields.
What does it cost?
Dataset queries are billed pay-on-success on credit-based plans (you're charged for results, not failed requests), with 2,000 free credits a month and no card to start. See pricing for plan limits.
Build on it
Put structured web data in your product.
Query 5.2M apps, Google Maps businesses and the live web indices through one REST API — clean JSON, no parsing, pay on success.