Direct answers to the questions buyers ask about web data. One question per page, every number sourced.
A 403 means the site refused your request, usually after a bot check. On 51 blocked pages, browser headers fixed 17 and a Chrome TLS fingerprint fixed 8 more.
Claude's bots do, including the one that fetches pages you ask about. ChatGPT's crawlers do, but OpenAI says robots.txt may not apply to ChatGPT-User.
Mostly by not triggering it. One plain request to 100 bot-protected pages drew 7 CAPTCHA puzzles, 29 silent checks and 19 flat blocks. What to do about each.
A full-site crawl is two jobs: find every URL, then fetch every page. On 38 of 77 sites we tested, a plain HTTP client could not get the sitemap file.
Find who sent the 429, pace each site, honor Retry-After and back off with jitter. In our Sep 16, 2026 run, 85% of rate-limit failures came from the API.
Most product pages already carry their own JSON. Read it first, then use a JSON Schema extraction API for the gaps. Tested on 100 sites, Oct 2026.
Run a trial on your own URLs, several times each, with a content check. In our Sep 16, 2026 run, 10-site trials ranked close providers wrong 40% of the time.
You can't scrape the past. Where history for a page can come from, how thin web archives are on the pages buyers need, and what to start collecting today.
Check each response, each record and each run, then sample by hand. In our Sep 16, 2026 benchmark run, 7.7% of 2xx responses lacked the page's content.
Test the body for content only the real page has. A marker, an item count and a size band catch most soft blocks, login walls and challenge pages.
Build when your targets are few, unprotected and stable. Buy when protected sites are in scope. A cost model to fill in, and our benchmark on the hard tail.
Akamai scores your TLS fingerprint, headers, a JavaScript sensor and behaviour. On 20 Akamai sites, 16 scraping APIs scored between 98% and 28%.
A scraping API fetches pages and you own the pipeline. A managed service delivers the dataset and owns upkeep. When each fits and what the SLA should cover.
Same code, different request: a cloud IP, a different TLS stack, and a headless container browser. How to tell which one is blocking you, and how to fix it.
A proxy sells IP addresses, an unblocker adds anti-bot handling, and a scraping API takes a URL and returns the page. When each fits, with benchmark data.
No. robots.txt is a voluntary protocol, not law, and no US court has held that ignoring it is by itself unlawful. It still changes your legal position.
hiQ won the CFAA fight and still lost the case. The Ninth Circuit protected public scraping; a contract claim ended the company. Both halves still apply.
Prices are facts, and facts are not copyrightable. Public price collection is the best-supported case in scraping law. The constraints are contract and antitrust.
Scraping public data is generally lawful in the US, and the real risk sits in contract and privacy law rather than hacking statutes. What the cases actually held.