25 step-by-step guides, grouped by what you are collecting. Every one shows the DIY Python route with requests and BeautifulSoup, why that route breaks in production, and the structured API call that returns the same data as JSON — plus the legal and anti-bot reality for that specific site.
Engine results, demand curves, local listings, and the academic index.
Independent-index web results without Google's bot defenses in the way.
Interest over time, rising and top queries — the official API is still allowlisted.
Places, categories, hours, and coordinates for local-business datasets.
Papers, citation counts, and authors from the literature index.
Catalogs, prices, sellers, and variants across the major marketplaces.
Ratings and review text, wherever customers actually leave it.
Business ratings and review text past the Places API's five-review cap.
Business profiles, category rankings, and full review streams.
Mobile ratings and reviews from both stores via the iTunes lookup pattern.
Hotels, attractions, and traveler reviews with locale-aware pagination.
Listings, availability, and pricing for places people stay or buy.
Catalogs, charts, and where a title is actually available.
Stores, player counts, and live match data.
Prices, fundamentals, and the market context around them.
The three guides that apply to every site on this page.
Pick the tool before you write the scraper — managed APIs, proxies, and what each is actually for.
The search-results side: which SERP APIs return clean rankings, and at what cost.
Public data, ToS, hiQ v. LinkedIn, the CFAA, and where personal data changes the answer.
Looking for a site that is not listed? The platform APIs cover more sources than there are guides, and the generic scraping API handles any URL you can point it at. New guides are added to this page as they publish — see the blog for everything else we write.
Every guide here ends at the same place: a documented endpoint that returns normalized JSON and bills only on success.
Social platforms
Profiles, posts, comments, and the engagement signals behind them.
Reddit
Subreddit posts, comment trees, and user history for sentiment and RAG pipelines.
YouTube
Videos, channels, comments, and transcripts — including the caption track.
TikTok
Posts, profiles, hashtag challenges, and the trending surfaces.
Instagram
Public profiles, posts, and reels without a logged-in session.
Twitter / X
Profiles and posts after the API pricing changes closed off the cheap routes.
LinkedIn
Company pages and products — and where the hiQ ruling actually leaves you.