Tony Wang7 min readHow to Scrape Shopify Stores in 2026 (API & Python)
Three ways to scrape Shopify store products and collections in 2026 — DIY Python, no-code tools, or a structured API — what each returns and the legal basics.
The fastest way to scrape Shopify stores in 2026 is to call a structured API that returns normalized JSON — products, variants, prices, collections, and store metadata — instead of crawling each storefront and parsing it yourself. DIY in Python is possible because most Shopify stores expose a public products feed, but variant handling, pagination, and per-store quirks make a maintained endpoint the simpler path at scale.
Shopify's own Admin and Storefront APIs require the store owner's credentials — they're for stores you control. To research storefronts you don't own, you collect the public product surface, which is what a structured scraping API normalizes for you.
Is it legal to scrape Shopify stores?
Storefront product pages are public, and collecting public data is generally treated differently from accessing private accounts — with the usual conditions:
- Collect only public storefront data — no admin or checkout access.
- Respect each store's terms and robots directives, plus your local law.
- Don't reuse product imagery or copy beyond what your use case and law allow.
- You are responsible for lawful, good-faith use of what you collect.
Not legal advice — see Is web scraping legal in 2026? for the full picture.
Option 1: DIY in Python (and why it breaks)
Many Shopify stores expose a public products.json feed, so a first pass looks easy:
import csv, requests
resp = requests.get(
"https://www.allbirds.com/products.json",
params={"limit": 250, "page": 1}, # walk page=1,2,... until the list is empty
headers={"User-Agent": "Mozilla/5.0"},
)
rows = []
for p in resp.json()["products"]:
for v in p["variants"]: # flatten nested variants into rows
rows.append({"title": p["title"], "handle": p["handle"],
"variant": v["title"], "price": v["price"], "available": v["available"]})
with open("shopify.csv", "w", newline="") as f:
w = csv.DictWriter(f, fieldnames=rows[0].keys()); w.writeheader(); w.writerows(rows)
# ...then handle HTTP 430 rate limits and stores that disable /products.json
Where it gets expensive:
- Inconsistent exposure — some stores disable the public feed, and Shopify returns HTTP
430when you hit its rate limit, so you need backoff, proxies, and fallbacks. - Variant flattening — each product has nested variants, options, and images to normalize into rows.
- Pagination —
/products.jsoncaps at 250 per page; large catalogs span many pages to walk and dedupe, and/collections.jsonis a second crawl. - Per-store differences — currencies, availability, and metafields vary across themes.
Option 2: No-code and ready-made tools
Point-and-click exporters can dump one store, but catalog and price monitoring means re-checking many stores on a schedule and storing history — a pipeline that an API serves better than a manual tool.
Option 3: A structured Shopify API
Crawlora's Shopify API wraps product, collection, and store endpoints behind one API key, returning normalized JSON. Point it at a store URL:
curl -G "https://api.crawlora.net/api/v1/shopify/products" \
-H "x-api-key: $CRAWLORA_API_KEY" \
--data-urlencode "url=https://www.allbirds.com" \
--data-urlencode "limit=50"
import requests
resp = requests.get(
"https://api.crawlora.net/api/v1/shopify/products",
headers={"x-api-key": "YOUR_API_KEY"},
params={"url": "https://www.allbirds.com", "limit": 50},
)
for product in resp.json()["data"]["products"]:
print(product["title"], product["price"], product["handle"])
A response is normalized JSON you can store directly (fields are illustrative — confirm the schema in the docs):
{
"code": 200,
"msg": "OK",
"data": {
"products": [
{
"handle": "wool-runner",
"title": "Wool Runner",
"price": 98.0,
"currency": "USD",
"available": true,
"variants": [{ "title": "US 9", "price": 98.0, "available": true }]
}
]
}
}
From there, map the catalog with collections, pull full product detail by handle, or read store metadata — all from the same key (every endpoint takes the store url):
h = {"x-api-key": "YOUR_API_KEY"}
base, store = "https://api.crawlora.net/api/v1/shopify", "https://www.allbirds.com"
collections = requests.get(f"{base}/collections", headers=h, params={"url": store}).json()["data"]
product = requests.get(f"{base}/products/wool-runner", headers=h, params={"url": store}).json()["data"]
meta = requests.get(f"{base}/store", headers=h, params={"url": store}).json()["data"]
Use /collections/{handle}/products to walk one collection, and the sitemap endpoints to discover every product and collection URL on the storefront — handy when /products.json is disabled.
A handful of popular Shopify-backed brands — Allbirds, Gymshark, Fashion Nova, and others — also have dedicated pinned-URL endpoints that return the same data without a url parameter at all, since the storefront address is fixed server-side.
Headless Shopify storefronts and facets
/products.json covers most stores, but not all of them. Some merchants run a headless storefront — a Next.js or React front end that renders pages server-side against a Shopify catalog backend but never exposes the classic feed or a discoverable *.myshopify.com domain at all. The same endpoints handle both kinds of store transparently: pass the same url, no extra flag. When the classic feed and the myshopify-domain fallback are both unavailable, the API falls back to parsing the storefront's own embedded search-result payload instead.
A live example: Gymshark's US storefront, www.gymshark.com, is a classic Shopify store — /products.json resolves normally. Its international storefront, row.gymshark.com, is headless: /products.json and /collections.json both fail, and there's no myshopify domain anywhere in the page. It's still a Shopify-backed catalog underneath, and the same call works unchanged:
curl -G "https://api.crawlora.net/api/v1/shopify/collections/leggings/products" \
-H "x-api-key: $CRAWLORA_API_KEY" \
--data-urlencode "url=https://row.gymshark.com" \
--data-urlencode "limit=3"
import requests
resp = requests.get(
"https://api.crawlora.net/api/v1/shopify/collections/leggings/products",
headers={"x-api-key": "YOUR_API_KEY"},
params={"url": "https://row.gymshark.com", "limit": 3},
)
data = resp.json()["data"]
print(data["transport_mode"], data["total_items"], data["total_pages"])
That call returns real listing and facet data for the collection (186 products across 4 pages at the time of writing), flagged with transport_mode: "ssr_embedded":
{
"code": 200,
"msg": "OK",
"data": {
"store_url": "https://row.gymshark.com",
"source_url": "https://row.gymshark.com",
"collection": "leggings",
"page": 1,
"limit": 3,
"total_items": 186,
"total_pages": 4,
"transport_mode": "ssr_embedded",
"products": [ /* 3 products, each shaped like the product below */ ],
"facets": {
"fit": { "regular": 128, "tall": 18, "short": 10 },
"canonicalColour": { "black": 90, "pink": 12 }
},
"facets_stats": {
"price": { "min": 27, "max": 85, "avg": 56.46 }
}
}
}
facets maps a filter field to its value-to-count buckets — the same data driving the storefront's own filter sidebar — and facets_stats gives min/max/avg for numeric fields like price. Facet field names mirror whatever the storefront itself uses (Gymshark's own fields are camelCase, like canonicalColour), so treat them as store-specific rather than a fixed list, and both are omitted entirely when a storefront's payload doesn't include facet data.
Products from a headless storefront also carry fields the classic feed never returns: colour, canonical_colour, discount_percentage, rating, rating_count, collection_tags, labels, and per-variant inventory_quantity. Here's a real one, pulled from row.gymshark.com's own product page:
{
"handle": "gymshark-train-t-shirt-ss-tops-black-aw26",
"title": "Train T-Shirt",
"price": 35,
"colour": "black",
"canonical_colour": "black",
"rating": 4.2,
"rating_count": 1549,
"collection_tags": ["all-products", "new-releases"],
"labels": ["new"],
"variants": [
{ "title": "XS", "price": 35, "available": true, "inventory_quantity": 16 }
]
}
No separate transport to configure — same base URL, same x-api-key header, same params as the classic path above.
The same fallback also covers three more endpoints beyond listings: product recommendations (intent=related or intent=complementary, same params as the classic path), static pages (parsed from the storefront's own CMS instead of a /pages.json feed), and the collections list (enumerated from the storefront's own sitemap). Same url, no flags:
curl -G "https://api.crawlora.net/api/v1/shopify/products/gymshark-train-t-shirt-ss-tops-black-aw26/recommendations" \
-H "x-api-key: $CRAWLORA_API_KEY" \
--data-urlencode "url=https://row.gymshark.com" \
--data-urlencode "intent=related"
One caveat on the collections list: without a /collections.json feed to read from, title is a best-effort guess from the URL slug (all-products → "All Products"), not the storefront's real display name, and there's no product count per collection — lower fidelity than the classic path. Treat it as discovery and confirm the real detail with /shopify/collections/{handle}/products.
Sorting and filtering results
On a headless storefront, /shopify/products and /shopify/collections/{handle}/products also accept a sortBy param and per-facet filter params, forwarded to the storefront's own search index exactly the way its sort/filter UI would send them. sortBy is one of sortLTH (price low to high), sortHTL (price high to low), or newest; filters are one query param per facet field name (the same names that show up in that listing's own facets field), comma-separated for multiple values within a facet and combined with & across facets:
curl -G "https://api.crawlora.net/api/v1/shopify/collections/leggings/products" \
-H "x-api-key: $CRAWLORA_API_KEY" \
--data-urlencode "url=https://row.gymshark.com" \
--data-urlencode "sortBy=sortLTH" \
--data-urlencode "fit=regular" \
--data-urlencode "limit=5"
Real numbers from row.gymshark.com's leggings collection (186 products unfiltered): sorting by sortLTH returns the first five at 21, 21, 27, 33, 33 (non-decreasing); sortHTL returns 85, 85, 85, 81, 76.5. Filtering by fit=regular alone narrows total_items from 186 to 128; stacking canonicalColour=black on top narrows it further to 34 — a real server-side narrowing, not client-side UI state. The response echoes back what was actually applied:
{
"total_items": 34,
"sort": "sortLTH",
"filters": { "fit": ["regular"], "canonicalColour": ["black"] }
}
Classic Shopify stores reject sortBy and filters outright, with a typed 400 error, instead of silently ignoring them. That's not a missing feature on our side — Shopify's own public /products.json feed has no server-side sort or filter at all. Requesting the same classic store's JSON with sort_by=price-ascending versus sort_by=price-descending returns byte-identical product order; sorting only happens in the theme's own page (or via the Storefront API, which needs the store's own credentials). So calling www.gymshark.com (a classic store) with sortBy set returns an error instead of a response that looks sorted but isn't:
{ "code": 400, "msg": "sort and filters are only supported for headless storefronts served via the embedded-SSR-JSON fallback transport (transport_mode ssr_embedded); Shopify's classic public catalog JSON has no server-side sort or filter support" }
If you need sort or filter behavior on a classic store, apply it client-side after fetching — the same thing the storefront's own theme does.
What you can collect
- Products: title, handle, price, availability, images, and variants
- Collections and the products within a collection, plus a lower-fidelity collection list on headless stores
- Store metadata, static pages, and storefront sitemaps
- Search suggestions and product recommendations
- On headless storefronts: filter facets and price/discount stats (
facets,facets_stats), and sorted/filtered listings (sortBy, per-facet filters)
Limitations and common challenges
- Not every store exposes the feed. Some disable
/products.jsonor rate-limit it (HTTP430), and some headless storefronts never expose it at all. The store, products, collection-products, product-recommendations, pages, and collections endpoints all fall back automatically to a public*.myshopify.comdomain or the storefront's own embedded payload — onlysearch/suggeststill depends on the classic feed being reachable; it has no fallback yet. Sitemaps aren't affected either way, since they're static files independent of the products feed. - Variants and metafields. Products nest variants, options, and images; flatten them into rows and expect theme-specific metafields to differ across stores.
- Pagination. Products and collections page at 250 max per page; walk and dedupe across pages.
- Imagery and copy are copyrighted. Prices and availability are facts you can collect; product photos and descriptions carry copyright — don't republish beyond what your use case and law allow.
Sources
Where this fits
Try it first, free: run any public URL through the Free Web Scraper, or check whether a site blocks bots with the Anti-Bot Checker — no signup.
Shopify data powers catalog monitoring, competitor price tracking, and assortment research. Combine it with how to scrape Shop.app for the consumer marketplace view — the cross-store surface where these same merchants get discovered — and the Amazon scraping API for cross-channel pricing, all under the e-commerce product intelligence workflow. For the same playbook on other marketplaces, see how to scrape Amazon product data and how to scrape eBay, or how to choose a web scraping API.
Get started by testing the endpoint in the Playground, reading the request and response schema in the API docs, and reviewing credit costs on the pricing page.
Part of our how-to-scrape guide series — every platform we cover, in one index.
Frequently asked questions
Can I scrape Shopify stores without getting blocked?
Crawlora handles proxy routing, pacing, retries, and fallbacks behind the API and returns normalized JSON, including for stores that rate-limit (HTTP 430) or disable the public products feed.
Does every Shopify store have a products.json feed?
Most do — /products.json is a credential-free endpoint returning up to 250 products per page — but some stores disable it or rate-limit it (HTTP 430). When the vanity domain blocks it, Crawlora can fall back to a public *.myshopify.com domain, and the sitemap endpoints discover product and collection URLs as a backup.
Doesn't Shopify already have an API?
Shopify's Admin and Storefront APIs require the store owner's credentials and are for stores you control. To research storefronts you don't own, Crawlora collects the public product surface as normalized JSON.
What data can I collect?
Products with variants, prices, availability, and images; collections and the products within them; store metadata; static pages; storefront sitemaps; search suggestions; and product recommendations. Prices and availability are facts; product imagery and copy are copyrighted.
How do I target a specific store?
Pass the store URL to the products endpoint. Use the product handle (with the store url) for full detail, the collections endpoint to map catalog structure, and the sitemap endpoints to enumerate URLs.
Can I monitor prices and catalog changes?
Yes. Re-run product and collection calls on a schedule and store the history to track price moves and assortment changes.
Can I scrape headless or JS-rendered Shopify stores, not just classic ones?
Yes. When a storefront's classic products.json feed and myshopify-domain fallback are both unavailable — common on headless Next.js/React storefronts — the store, products, collection-products, product-recommendations, pages, and collections endpoints all fall back automatically to parsing the storefront's own embedded payload, with no extra flag needed. Responses from this path carry transport_mode: "ssr_embedded" plus richer per-product fields (colour, discount percentage, rating, collection tags, per-variant inventory quantity) and, on the products and collection-products endpoints, filter facets with live counts and price/discount stats. Only search/suggest doesn't have this fallback yet.
Can I sort or filter product listings?
Yes, on headless storefronts served via the ssr_embedded transport. /shopify/products and /shopify/collections/{handle}/products both accept a sortBy param (sortLTH, sortHTL, or newest) and dynamic facet-filter params, one query param per facet field (for example fit=regular, comma-separated for multiple values within a facet, combinable across facets with &). Classic Shopify stores reject both with a 400 invalid-param error instead of silently ignoring them, because Shopify's own public /products.json feed has no server-side sort or filter support at all.