How to scrape Amazon product data in 2026
Last updated: September 29, 2026. Benchmark figures come from the September 16, 2026 run of the Web Data Frontier Benchmark. Every script on this page was run on September 29, 2026, and the output shown is what it printed.
To scrape Amazon product data in 2026, fetch the product page at https://www.amazon.com/dp/<ASIN> from a US address and read the title, buy-box price, rating, review count and availability out of the HTML by their element ids. Search pages work the same way, one s-search-result block per product. Getting a page back is the easy part. On September 29, 2026, a plain Python request returned product pages without trouble, and Amazon's search page answered it with HTTP 503 and its automated-access notice. The harder problem is the price. From a connection outside the US, 29 of 32 product pages came back HTTP 200 with the price in Moroccan dirham, because Amazon prices each page for the delivery location it guesses from your IP. The Python below returned three products with US dollar prices through the String Web Access API, and the TypeScript returned a search page with 16 results.
This page is part of our comparison of the best e-commerce scraping APIs.
TL;DR
- A plain
requests.geton a product page returned HTTP 200 and the full 1.5 MB page on three of three tries. The same call on a search page got HTTP 503 and Amazon's "To discuss automated access to Amazon data" notice every time. - A Chrome User-Agent was enough to get search results: 16 to 22 per page.
- The failure that costs money is a correct-looking page with the wrong price. Of 32 product pages fetched from a non-US connection, 32 returned a title, 30 carried a price, and 29 of those prices were in Moroccan dirham. Route requests through the country you want prices for.
- In the September 16, 2026 benchmark, 9 of 16 scraping APIs returned amazon.com on all five attempts. Oxylabs and Decodo returned nothing, each with an HTTP 400 from its own API.
- Every Amazon request we logged through String billed as a premium fetch: $3.00 per 1,000 on Starter and $2.00 on Growth.
What people who scrape Amazon run into
The complaints are old and they have not changed much. In 2021, a maintainer of the unfurl.js link-preview library was asked why Amazon URLs sometimes came back titled "Sorry! Something went wrong!". The reporter later found that Amazon had returned HTTP 503 with that page, and asked how to detect it so the bad metadata never reached a database (GitHub, jacktuck/unfurl issue #77, May 30, 2021).
In 2023 a Stack Overflow user described the next stage: requests failed "even when using cookie and headers", while the same request from cURL returned the page (Stack Overflow, MarinaF, August 25, 2023). That is the same split we saw in 2026: the block depends on what the request looks like and where it comes from, and it differs between product pages and search pages.
Why a plain request to Amazon fails, sometimes
We sent three kinds of request to one product page and one search page on September 29, 2026, three times each. The product page is the benchmark's Amazon URL, a bike lock with ASIN B08KCWFMRS.
| What we tried, September 29, 2026 | Product page | Search page (/s?k=wireless+earbuds) |
|---|---|---|
requests.get, default User-Agent |
HTTP 200, 1.5 MB, title and price, 3 of 3 | HTTP 503, 2.7 KB, automated-access notice, 3 of 3 |
requests.get, Chrome User-Agent |
HTTP 200, 2.8 to 3.1 MB, title and price, 3 of 3 | HTTP 200, 16 to 22 results, 3 of 3 |
curl_cffi, impersonating Chrome |
HTTP 200, 2.4 to 2.9 MB, title and price, 3 of 3 | HTTP 200, 16 to 22 results, 3 of 3 |
Then we fetched 32 product pages back to back with a Chrome User-Agent: 16 ASINs, twice each, in 69 seconds. All 32 came back HTTP 200 with a product title. None hit a Robot Check. By status code and by title, that is a clean run.
The prices say otherwise. The connection was outside the US, and the page's "Deliver to" line read Morocco on all 32. Thirty pages carried a price, and 29 of them were in dirham: MAD208.56 for the bike lock that costs $21.59 from a US address. Two pages had an empty price field. A scraper that checks for a 200 and a title would store all of it.
So a plain request is enough to get pages from a small, occasional job, and the thing to fix first is the location. The two other reasons people move to a service are volume, where Amazon's Robot Check and 503s start, and search pages, which block a default client outright.
Ways to get Amazon product data, compared
| Method | Product pages | Search results | What you maintain |
|---|---|---|---|
Plain requests |
Yes, at low volume | No, HTTP 503 | Headers, a US exit, and detection for the 503 page |
requests or curl_cffi with a browser profile |
Yes, at low volume | Yes, at low volume | The above, plus impersonation profiles |
| Headless browser (Playwright, Selenium) | Yes, slower | Yes, slower | A browser fleet and proxies |
| Amazon's Product Advertising API | Yes, for Amazon Associates | Yes, for Amazon Associates | Associates approval and Amazon's usage terms |
| A prebuilt scraper, such as String's Amazon Actors on Apify | Yes, by ASIN | Yes, first page per keyword | An Apify account |
| A fetch API such as the String Web Access API | Yes | Yes | Your parser, as below |
Sellers can also read their own catalog and orders through Amazon's Selling Partner API. That covers your own listings, not the rest of the store.
What 16 scraping APIs did with amazon.com
The Web Data Frontier Benchmark sends the same 100 URLs to 16 scraping APIs, five attempts each. The Amazon URL is the bike lock product page above. The harness and targets are open source, and so is the raw result file for the September 16 run.
| API | amazon.com, 5 attempts | Overall success, 100 sites |
|---|---|---|
| String | 100% | 97.0% |
| Scrapfly | 100% | 86.2% |
| ScraperAPI | 100% | 84.0% |
| Firecrawl | 100% | 80.2% |
| Bright Data | 100% | 74.6% |
| ScrapingBee | 100% | 73.0% |
| Context.dev | 100% | 72.0% |
| Nimble | 100% | 68.6% |
| Zyte | 100% | 68.0% |
| ZenRows | 80% | 41.2% |
| ScrapingAnt | 80% | 36.4% |
| Browserbase | 60% | 41.4% |
| Apify | 40% | 77.4% |
| Scrapingdog | 40% | 45.6% |
| Oxylabs | 0% | 69.0% |
| Decodo | 0% | 50.6% |
Nine of 16 APIs returned amazon.com on all five attempts: 60 passes out of 80 requests. The average API passed 75% of attempts on Amazon, against 66.6% across all 100 sites and 62.1% across the 17 retail and e-commerce sites. Amazon ranked 63rd of 100 by difficulty, which puts it in the middle of the field. String averaged 2.9 seconds per Amazon attempt.
The pass check on this target is strict. The benchmark counts a pass only when the body contains the product's full title, "Bike Lock Heavy Duty Anti Theft, Keyed Bike U Lock with 4FT Security Cable and Mounting Bracket for Road Bike, Mountain Bike, Folding Bike". A Robot Check page or the 503 notice cannot pass it. Most partial scores in the raw file say "Missing expected text": the API returned a page, and the page was not the product.
The two zeros need context. Oxylabs returned HTTP 400 from its own API on all five attempts in about 0.6 seconds, and Decodo returned HTTP 400 in about 0.2 seconds. Neither reached Amazon. The benchmark calls each vendor's general-purpose endpoint, and Oxylabs sells separate Amazon-specific sources that this run did not test. Our String vs Oxylabs page covers that difference. The benchmark also does not check which currency a page came back in, so a pass here says the product page loaded, not that the price is the one you want.
How to scrape Amazon with Python
Install requests, set your key, and pass ASINs on the command line. The script asks String for a US exit with countryCode and prints the "Deliver to" line for every product. Amazon prices follow the delivery location, so fetch again any page whose location is not the one you asked for.
"""Scrape Amazon product pages by ASIN with the String Web Access API.
pip install requests
export STRING_API_KEY=...
python amazon_products.py B08KCWFMRS B0CRTR3PMF
"""
import html as htmllib
import os
import re
import sys
import requests
API = "https://request.usestring.ai/v1/fetch"
def fetch(url: str) -> str:
response = requests.post(
API,
headers={"Authorization": f"Bearer {os.environ['STRING_API_KEY']}"},
# countryCode US: Amazon prices and availability follow the delivery location.
json={"url": url, "countryCode": "US"},
timeout=90,
)
response.raise_for_status()
page = response.json()
if page["statusCode"] != 200:
raise RuntimeError(f"amazon.com returned {page['statusCode']} for {url}")
return page["data"]
def first(pattern: str, text: str):
m = re.search(pattern, text, re.S)
return htmllib.unescape(re.sub(r"\s+", " ", m.group(1))).strip() if m else None
def product(asin: str) -> dict:
page = fetch(f"https://www.amazon.com/dp/{asin}")
if "validateCaptcha" in page or "api-services-support@amazon.com" in page:
raise RuntimeError(f"{asin}: Amazon served a Robot Check page")
if 'id="productTitle"' not in page:
raise RuntimeError(f"{asin}: HTTP 200 but no product title: treat this as blocked")
return {
"asin": asin,
"title": first(r'id="productTitle"[^>]*>(.*?)</span>', page),
# The buy-box price. None means the page loaded without one: check deliver_to.
"price": first(r'id="corePrice(?:Display_desktop)?_feature_div".*?class="a-offscreen">([^<]+)<', page),
"rating": first(r'id="acrPopover"[^>]*title="([\d.]+) out of 5 stars', page),
"reviews": first(r'id="acrCustomerReviewText"[^>]*>([^<]+)<', page),
"availability": first(r'id="availability".*?<span[^>]*>(.*?)</span>', page),
"deliver_to": first(r'id="glow-ingress-line2"[^>]*>(.*?)<', page),
}
if __name__ == "__main__":
asins = sys.argv[1:] or ["B08KCWFMRS", "B0CRTR3PMF", "B09FT58QQP"]
for asin in asins:
p = product(asin)
print(f'{p["asin"]} {p["price"]} | {p["rating"]} stars, {p["reviews"]} | {p["availability"]} | deliver to {p["deliver_to"]}')
print(f' {(p["title"] or "")[:80]}')
What it printed on September 29, 2026, in 17.7 seconds for the three products:
B08KCWFMRS $21.59 | 4.5 stars, (3,278) | In Stock | deliver to Update location
Bike Lock Heavy Duty Anti Theft, Keyed Bike U Lock with 4FT Security Cable and M
B0CRTR3PMF $24.99 | 4.3 stars, (42,977) | In Stock | deliver to Update location
Soundcore P30i by Anker Noise Cancelling Earbuds, Hands-Free Viewing
B09FT58QQP $14.99 | 4.3 stars, (117,603) | In Stock | deliver to Update location
TOZO A1 Wireless Earbuds Bluetooth 5.3 Light Weight in Ear Headphones
"Update location" is the line Amazon shows when no ZIP code is set, and the prices are in US dollars. If you need the price for a specific ZIP, which can differ for groceries and some sellers, you need a session with that ZIP set, and this script does not do that.
How to scrape Amazon search results with TypeScript
No dependencies. Node 22.6 or later runs TypeScript directly (we ran it on Node 24), and so does Bun.
// Scrape one Amazon search results page with the String Web Access API.
// export STRING_API_KEY=...
// node amazon_search.ts "wireless earbuds" (Node 22.6+; or: bun amazon_search.ts "...")
type FetchEnvelope = { statusCode: number; data: string };
async function fetchPage(url: string): Promise<string> {
const response = await fetch("https://request.usestring.ai/v1/fetch", {
method: "POST",
headers: { Authorization: `Bearer ${process.env.STRING_API_KEY}`, "Content-Type": "application/json" },
body: JSON.stringify({ url, countryCode: "US" }),
});
if (!response.ok) throw new Error(`String returned ${response.status}: ${await response.text()}`);
const page = (await response.json()) as FetchEnvelope;
if (page.statusCode !== 200) throw new Error(`amazon.com returned ${page.statusCode}`);
return page.data;
}
const query = process.argv[2] ?? "wireless earbuds";
const html = await fetchPage(`https://www.amazon.com/s?k=${encodeURIComponent(query)}`);
if (html.includes("api-services-support@amazon.com") || html.includes("validateCaptcha")) {
throw new Error("Amazon served its automated-access page: treat this as blocked");
}
// Each result is a <div data-asin="..." data-component-type="s-search-result"> block.
const starts = [...html.matchAll(/<div[^>]*data-asin="([A-Z0-9]{10})"[^>]*data-component-type="s-search-result"/g)];
const decode = (s: string) => s.replace(/&/g, "&").replace(/'/g, "'").replace(/"/g, '"').trim();
const results = starts.map((m, i) => {
const block = html.slice(m.index, starts[i + 1]?.index ?? m.index + 60000);
const text = (re: RegExp) => {
const hit = block.match(re)?.[1];
return hit ? decode(hit) : undefined;
};
return {
asin: m[1],
title: text(/<h2[^>]*>[\s\S]*?<span[^>]*>([^<]+)<\/span>/),
price: text(/class="a-offscreen">([^<]+)</),
rating: text(/([\d.]+) out of 5 stars/),
sponsored: /Sponsored/.test(block.slice(0, 4000)),
};
});
if (results.length === 0) throw new Error("HTTP 200 but no results in the page: treat this as blocked");
console.log(`"${query}": ${results.length} results, ${results.filter((r) => r.price).length} with a price`);
for (const r of results.slice(0, 5)) {
console.log(`${r.asin} ${r.price ?? "no price"} ${r.rating ?? "-"} stars${r.sponsored ? " (sponsored)" : ""} ${(r.title ?? "").slice(0, 60)}`);
}
We ran it on September 29, 2026, on Node 24, in 4.4 seconds:
"wireless earbuds": 16 results, 14 with a price
B0H5HZZ8S1 $20.98 5.0 stars Wireless Earbuds, Sports Bluetooth Headphones, Running Headp
B0FQFB8FMG $234.00 4.4 stars Apple AirPods Pro 3 Wireless Earbuds with Active Noise Cance
B0CRTR3PMF $24.99 4.3 stars Soundcore P30i by Anker Noise Cancelling Earbuds, Hands-Free
B0HDB41H28 $19.98 5.0 stars Wireless Earbuds, Bluetooth 5.3 Headphones HiFi Stereo 50H P
B09FT58QQP $14.99 4.3 stars TOZO A1 Wireless Earbuds Bluetooth 5.3 Light Weight in Ear H
A run on Bun a few seconds later returned the same 16-result page with a different order. Amazon reorders search results between loads, so do not treat rank from one fetch as stable. Two of the 16 results carried no price. Take the ASINs from this page and pass them to the Python script for the full product record.
The failure mode: HTTP 200 with the wrong price
Amazon's failures come in three shapes, and only one of them has an error code.
- The 503 notice, 2.7 KB. What a default client got on search pages. It says "To discuss automated access to Amazon data please contact api-services-support@amazon.com". Some older pages title it "Sorry! Something went wrong!", which is what the unfurl.js report above ran into.
- The Robot Check. A captcha page with a
validateCaptchaform. We did not hit it in our 32-page burst, and it is what people report at higher volume. Both scripts check for it. - A real page priced for the wrong place. HTTP 200, full title, a price in a currency or at a level you did not expect. This is the one to plan for, because it passes every check that looks at status or structure.
Two checks catch the third one. Read the "Deliver to" line (glow-ingress-line2) on every product page and reject any that is not the location you asked for. And check the currency symbol on every price before you store it. The Python script prints both. A request that returns a page is billed whether or not the price is the right one, so catch it before it reaches your database.
What it costs
Both Amazon requests whose billing header we logged on September 29, 2026, one product page and one search page, billed as a request-based fetch on a premium proxy. From String's pricing, that is $3.00 per 1,000 requests on Starter ($20 a month) and $2.00 per 1,000 on Growth ($100 a month). A request that fails is not billed.
Worked through from the runs above:
- Price monitoring. One request per ASIN per check. 10,000 ASINs checked once a day is 300,000 requests a month: $600 in usage on Growth, or $900 on Starter.
- Search tracking. One request per keyword per page. 1,000 keywords a day, first page only, is 30,000 requests a month: $60 on Growth, $90 on Starter.
If you only need a few hundred product pages now and then, a plain request from a US connection may be all you need. Test that first.
FAQ
How to scrape Amazon product data?
Fetch https://www.amazon.com/dp/<ASIN> from a US address and read the fields by id: productTitle for the title, the a-offscreen span inside corePrice_feature_div for the buy-box price, acrPopover for the rating, acrCustomerReviewText for the review count and availability for stock. Check the "Deliver to" line before you trust the price.
How do I scrape Amazon with Python?
Send the product URL to a fetch API with requests and parse the HTML with regular expressions or BeautifulSoup. The Python script on this page returned three products with title, US price, rating, review count and availability in 17.7 seconds on September 29, 2026.
Why does Amazon return a 503 error to my scraper?
A 503 with a small page that says "To discuss automated access to Amazon data please contact api-services-support@amazon.com" is Amazon's block response. In our test on September 29, 2026, a default Python requests call got it on every search page and never on the product page. A Chrome User-Agent was enough to get search results at low volume.
Why does my Amazon scraper return prices in the wrong currency?
Amazon prices a page for the delivery location it infers from your IP address. From a connection outside the US, 29 of the 30 prices we got on 32 amazon.com product pages were in Moroccan dirham. Route requests through a US exit, and check the "Deliver to" line and the currency symbol before storing a price.
Does Amazon have an API for product data?
Amazon's Product Advertising API returns product data to members of the Amazon Associates program, under Amazon's terms. Sellers can read their own listings through the Selling Partner API. Neither is a general way to pull any product's page, which is why most price monitoring scrapes the public product page.
How do I scrape Amazon search results?
Fetch https://www.amazon.com/s?k=<keyword> and split the page on the data-component-type="s-search-result" blocks. Each carries a data-asin, a title in its h2, a price in a-offscreen and a rating. The TypeScript script on this page returned 16 results for "wireless earbuds". Result order changes between loads.
How many Amazon pages can I scrape before getting blocked?
We fetched 32 product pages in 69 seconds from one home connection with a Chrome User-Agent and got no Robot Check. We did not push further. Reports of captchas and 503s come from higher volumes and from cloud servers, so plan for them if your job runs on a schedule.
Sources
- String benchmark: Web Data Frontier Benchmark, harness and targets, September 16, 2026 raw results, Amazon target definition
- String docs: Fetch API reference, including countryCode, pricing
- Amazon: Product Advertising API 5.0 documentation
- Oxylabs: Amazon targets documentation
- Practitioner reports: GitHub, jacktuck/unfurl issue #77, 2021, Stack Overflow, 2023
- Libraries: curl_cffi
- Related on usestring.ai: best e-commerce scraping APIs, how to scrape Zillow, String's Amazon Actors on Apify, String vs Oxylabs, the Web Access API, how to scrape Reddit, how to scrape Google search results, how to scrape X (Twitter)
Cheers,
String team
