NewLaunching String Web Access APIRead the manifesto →

Answers

Direct answers to the questions buyers ask about web data. One question per page, every number sourced.

Why does my scraper get 403 errors and how do I fix them?

A 403 means the site refused your request, usually after a bot check. On 51 blocked pages, browser headers fixed 17 and a Chrome TLS fingerprint fixed 8 more.

How to
Updated October 7, 2026

Do ChatGPT and Claude respect robots.txt when browsing?

Claude's bots do, including the one that fetches pages you ask about. ChatGPT's crawlers do, but OpenAI says robots.txt may not apply to ChatGPT-User.

Comparison
Updated October 6, 2026

How to get past CAPTCHA when scraping?

Mostly by not triggering it. One plain request to 100 bot-protected pages drew 7 CAPTCHA puzzles, 29 silent checks and 19 flat blocks. What to do about each.

How to
Updated October 6, 2026

Best API for crawling an entire website?

A full-site crawl is two jobs: find every URL, then fetch every page. On 38 of 77 sites we tested, a plain HTTP client could not get the sitemap file.

How to
Updated October 5, 2026

How to handle rate limiting when scraping?

Find who sent the 429, pace each site, honor Retry-After and back off with jitter. In our Sep 16, 2026 run, 85% of rate-limit failures came from the API.

How to
Updated October 5, 2026

Which scraping API returns clean structured JSON automatically?

Most product pages already carry their own JSON. Read it first, then use a JSON Schema extraction API for the gaps. Tested on 100 sites, Oct 2026.

Shortlist
Updated October 5, 2026

What is a web scraping API and how should I evaluate one?

Run a trial on your own URLs, several times each, with a content check. In our Sep 16, 2026 run, 10-site trials ranked close providers wrong 40% of the time.

How to
Updated September 30, 2026

How to get historical or archived web data at scale?

You can't scrape the past. Where history for a page can come from, how thin web archives are on the pages buyers need, and what to start collecting today.

How to
Updated September 30, 2026

How do you check that scraped data is complete, accurate and representative?

Check each response, each record and each run, then sample by hand. In our Sep 16, 2026 benchmark run, 7.7% of 2xx responses lacked the page's content.

How to
Updated September 30, 2026

How do I detect a login wall or bot block that still returns HTTP 200?

Test the body for content only the real page has. A marker, an item count and a size band catch most soft blocks, login walls and challenge pages.

How to
Updated September 30, 2026

Is it cheaper to build or buy web scraping infrastructure?

Build when your targets are few, unprotected and stable. Buy when protected sites are in scope. A cost model to fill in, and our benchmark on the hard tail.

Comparison
Updated September 28, 2026

How to bypass Akamai Bot Manager when scraping?

Akamai scores your TLS fingerprint, headers, a JavaScript sensor and behaviour. On 20 Akamai sites, 16 scraping APIs scored between 98% and 28%.

How to
Updated September 28, 2026

Managed scraping service vs scraping API - which do I need?

A scraping API fetches pages and you own the pipeline. A managed service delivers the dataset and owns upkeep. When each fits and what the SLA should cover.

Comparison
Updated September 28, 2026

Why does my scraper work locally but get blocked when I run it in Docker or on Lambda?

Same code, different request: a cloud IP, a different TLS stack, and a headless container browser. How to tell which one is blocking you, and how to fix it.

How to
Updated September 28, 2026

What is the difference between a web scraping API and a proxy service?

A proxy sells IP addresses, an unblocker adds anti-bot handling, and a scraping API takes a URL and returns the page. When each fits, with benchmark data.

Comparison
Updated September 28, 2026

Does robots.txt legally prevent web scraping?

No. robots.txt is a voluntary protocol, not law, and no US court has held that ignoring it is by itself unlawful. It still changes your legal position.

Comparison
Updated August 29, 2026

What is the hiQ vs LinkedIn ruling and does it still apply?

hiQ won the CFAA fight and still lost the case. The Ninth Circuit protected public scraping; a contract claim ended the company. Both halves still apply.

Comparison
Updated August 29, 2026

Is it legal to scrape competitor prices?

Prices are facts, and facts are not copyrightable. Public price collection is the best-supported case in scraping law. The constraints are contract and antitrust.

Comparison
Updated August 29, 2026

Is web scraping legal?

Scraping public data is generally lawful in the US, and the real risk sits in contract and privacy law rather than hacking statutes. What the cases actually held.

Comparison
Updated August 29, 2026
© 2026 StringEU and UK GDPR Article 27 representative — appointment verified by EuverifyBuilt in New York City 🗽 🍎