How to scrape a website without getting blocked in 2026
Last updated: October 5, 2026. Benchmark figures come from the September 16, 2026 run of the Web Data Frontier Benchmark.
To scrape without getting blocked in 2026, make every request look like one consistent, unhurried browser session. Match your TLS and HTTP/2 fingerprint to the browser your User-Agent claims (curl_cffi in Python), egress from residential or mobile IPs and keep one IP per session, cap concurrency per domain at one or two with jittered delays, hold cookies and a plausible Referer chain, and check the response body for the content you expected rather than trusting the status code, because most 2026 blocks arrive as a 200. Add a stealth browser (nodriver or Camoufox) only for pages that need JavaScript. Every one of those techniques has a date on it, and this guide names the thing that breaks each one. When the fixes turn into a standing job, a fetch API that bills per successful page is cheaper than the engineer. In the September 16, 2026 run of the open Web Data Frontier Benchmark (100 bot-protected sites, 16 scraping APIs), success rates ran from 36.4% to 97.0%; String's Web Access API was the 97.0%.
This page is part of our series on the best web scraping APIs, tested.
TL;DR
- Four signals stack: rate, IP reputation, connection fingerprint (TLS, HTTP/2, headers), and behavior. Trip three and you are gone.
- Most blocks in 2026 are soft: a 200 with a challenge page, an empty shell, or a real-looking page with zero results. Check for a content marker.
- The highest-leverage fix most people skip is the TLS and HTTP/2 fingerprint. A residential IP with a Python handshake is still a bot.
- Rotate per session, not per request. Pair each cookie jar with one IP.
- Concurrency is a fingerprint. On one Akamai-protected target, moving 30 chained requests from serial to concurrent took denials from 13% to 50% with nothing else changed.
- Fingerprint identity is scored per site. A newer Chrome TLS profile can be burned on a site where an older one passes.
- Cloudflare's Precursor (July 2026) and DataDome's per-customer models score whole sessions. Stateless clients hit a ceiling.
- Managed option: String, 97.0% in the September 16, 2026 benchmark run, billed per successful request.
Why sites block you
Four signals do most of the work, and they stack.
Rate is the crude one. Too many requests from one source in too short a window and you get a 429, a slowdown, or a silent block. Easy to trigger, easy to fix, and the least of your problems on a serious anti-bot site.
IP reputation is next. Datacenter ranges (AWS, GCP, Hetzner, the cloud you run from) are known and scored down. Residential and mobile IPs score neutral to positive because real users sit behind them. This is why proxies exist as a business.
Fingerprints are where it gets hard. Every HTTPS request starts with a TLS handshake, and the shape of that handshake (cipher order, extensions, the JA3 and now JA4 hashes derived from it) identifies the client library. Python requests and real Chrome do not produce the same handshake. Cloudflare, DataDome, Akamai, Kasada, and HUMAN (PerimeterX) all score TLS and HTTP/2 fingerprints against what the request claims in its User-Agent, and Cloudflare's docs list the mismatch as a primary signal (Cloudflare docs). Claim Chrome, hand over Python, and you have told on yourself before the server reads a header.
Behavior is the newest and hardest. Session paths, request timing, whether you touch the dynamic elements a human would, how the pointer moves. Cloudflare's Precursor, announced July 13, 2026, injects a session-scoped script and cross-checks pointer, keyboard, focus, and visibility events at the edge; a refresh does not reset the score. DataDome advertises tens of thousands of per-customer models, so the "normal" it compares you to is that site's own traffic (Scrapfly, updated Aug 2026). This is the signal a better HTTP client cannot fix, and the reason the rest of this guide has a ceiling.
First, learn to see the block
Most people debug the wrong thing because they never saw the block. In 2026 a hard 403 is the easy case. The common cases look like success.
A 200 with a challenge page in the body. Cloudflare's "Just a moment..." interstitial, a HUMAN "Press & Hold" page, or a Turnstile widget served from the site's own origin all arrive as 200s. If your pipeline checks response.ok, it records the block as a win and stores garbage.
A 200 with an empty shell. Anti-bot scripts often ship the page chrome and withhold the data until a sensor call succeeds. You get a title, a nav bar, and no product.
A 200 with a real page and zero results. This one is nasty. On one large travel site we hit a search page that returned 1.46MB and 5,620 listings under one Chrome TLS profile, and 1.23MB and zero listings under a newer Chrome profile, same URL, same proxy, same session, both HTTP 200 and both logged as success by the site's own telemetry. The site's bot SDK had decided off the handshake, and the empty result was the block. The newer fingerprint had been burned on that one target; the older one had not.
The rule that follows: define a content marker per target (a product name, a result count, a string that only the real page contains) and treat its absence as a block regardless of status. That is how the Web Data Frontier Benchmark scores a pass: 2xx plus the marker, so a challenge page at 200 fails. For a longer checklist, see how to detect a bot block that returns HTTP 200.
One more trap. Your own block detector can lie. We once burned two investigation passes on a "Turnstile challenge" that did not exist: the detector matched a turnstile test id that the target's page scaffolding ships on every healthy page too, and the real problem was a decommissioned listing returning a not-found route. One control fetch of a known-good page would have shown the marker present there as well. Before you buy proxies to beat an anti-bot error, fetch a page you know works and confirm the marker is absent on it.
The techniques, and when each one breaks
1. Respect robots.txt, rate limits, and concurrency
Start here because most blocks are self-inflicted. Read robots.txt, honor Crawl-delay, and cap concurrency per domain. A conservative default: one request in flight per domain per IP, a few seconds between requests, with jitter so you are not a metronome.
import time, random
from urllib.robotparser import RobotFileParser
rp = RobotFileParser()
rp.set_url("https://example.com/robots.txt")
rp.read()
if rp.can_fetch("MyBot/1.0", url):
fetch(url)
time.sleep(random.uniform(2, 5)) # jitter, not a fixed interval
Concurrency is itself a fingerprint. On one Akamai-protected car-inventory site, a chain of 30 requests per vehicle went from 13% denied to 50% denied when the chain was issued concurrently instead of serially, with the same proxies and the same fingerprint. Humans do not open thirty tabs at once. Serialize per session and parallelize across sessions.
Works until: the site's tolerance is lower than you guessed, or the block is fingerprint-based and has nothing to do with rate. Pacing keeps you off the easy blocklists. It does nothing against a fingerprint mismatch.
2. Rotate residential proxies, per session
Move off datacenter IPs. Residential and mobile proxies route through real consumer connections, so the reputation signal flips from bad to neutral. Proxy vendors sell residential bandwidth per gigabyte, and the rate falls with volume. Rotate per session rather than per request on any site that sets cookies, and hold a sticky IP through a multi-step flow. A visitor whose IP changes every request has no coherent session score, and on Cloudflare the __cf_bm cookie is exactly that score.
Two things proxies do not fix. A clean residential IP with a Python TLS handshake still looks like a bot. And if you drive a browser through an HTTP or SOCKS proxy, WebRTC can leak; some WAFs now require a live UDP WebRTC round trip that those proxies cannot carry (camoufox#672). Disable WebRTC or route it. Proxies also cost real money on rendered pages, since images and scripts burn far more bandwidth than raw HTML.
Works until: the fingerprint gives you away anyway.
3. Match your TLS, HTTP/2, and headers to one real browser
This is the highest-leverage fix most people skip. Use a client that impersonates a browser's handshake. In Python, curl_cffi wraps curl-impersonate and reproduces a real Chrome, Firefox, or Safari TLS and HTTP/2 fingerprint with the ergonomics of requests.
from curl_cffi import requests
r = requests.get(
"https://example.com",
impersonate="chrome", # newest Chrome profile the library ships
proxies={"https": "http://user:pass@residential-proxy:port"},
)
Three caveats that catch people. First, JA3 alone is no longer the target; Chrome began shuffling TLS extension order in 2023, and Cloudflare, Akamai, DataDome, and HUMAN now score JA4 and the HTTP/2 shape together. Second, versions drift. Chrome ships a new stable major version about every four weeks, and curl_cffi's browser profiles trail it. Use the impersonate="chrome" alias so you get the newest profile the library has, and make your User-Agent and sec-ch-ua client hints name the same major version as the TLS profile. Third, the whole header set matters: order, Accept-Language, Accept-Encoding, the sec-fetch-* values, and no leftover python-requests anything. A perfect handshake under a requests-shaped header block is still a mismatch. Verify at tls.peet.ws or ja4db.com against a real Chrome run from the same proxy.
Works until: the site requires JavaScript, or its behavioral engine wants a session a stateless client cannot fake. Also, per the travel-site example above, a profile can be burned on a specific site. If a target flips to empty results, try the previous major version's profile before you blame the proxy. For Cloudflare specifically, with scripts we ran on eight Cloudflare sites, see how to bypass Cloudflare when scraping.
4. Use a stealth browser when you need JavaScript
When the content only exists after JS runs, you need a real engine. The 2026 shift is away from patching a detectable browser and toward engines that were never detectable. puppeteer-extra-plugin-stealth applies around 29 patches, and those patches produce their own recognizable fingerprint; CreepJS classifies it as a stealth tool. nodriver (and its active fork zendriver) talks raw Chrome DevTools Protocol with no WebDriver layer. Camoufox modifies Firefox itself so the usual leaks never surface. In Ian Paterson's May 2026 test (31 targets, not all of them Cloudflare, one residential IP), nodriver passed 28 of 31, Patchright and Camoufox 25, and plain Playwright 24 (dev.to).
Works until: CDP itself gives you away. Cloudflare's docs list Chrome DevTools Protocol detection as a signal, and CDP artifacts survive deleting navigator.webdriver. A full browser is also far heavier than an HTTP request, so this does not scale cheaply.
5. Retry with backoff, and separate transient from durable
Blocks are often probabilistic. A request that fails now may pass in thirty seconds from a different IP. Retry with exponential backoff and jitter, and treat a soft block (a 200 without the marker) as a failure to retry.
import time, random
def fetch_with_backoff(url, attempts=5):
for i in range(attempts):
r = get(url)
if r.ok and has_marker(r.text):
return r
time.sleep((2 ** i) + random.uniform(0, 1)) # 1s, 2s, 4s, 8s...
raise BlockedError(url)
Keep two ledgers. Timeouts, 502s, and connection resets are usually transient and clear on retry. A challenge page, a login wall, or an empty shell that repeats from three IPs is durable, and retrying it only burns bandwidth and reputation. In one of our own test runs, three first-pass misses out of 35 targets recovered on a plain retry; the two that did not were real anti-bot walls. Score them differently.
Works until: the site blocks the identity durably. Backoff buys retries, not a way through.
6. Cache aggressively
The request you do not make cannot be blocked. Cache responses, dedupe URLs, and use conditional requests (If-Modified-Since, ETag) so you refetch only what changed. On a crawl that revisits pages, this cuts volume by an order of magnitude and directly lowers how often you trip rate and reputation signals.
Works until: you need fresh data. Caching reduces exposure; it does not unblock anything.
7. Manage sessions like a real user
On sites that track state, behave like one visitor across a session rather than a swarm of stateless requests. One cookie jar per proxy IP. Warm the session with a homepage or category page before the deep URL. Set a plausible Referer chain; a user who lands on a product page usually came from a category. Hold the cookies an anti-bot sets (__cf_bm, cf_clearance, datadome, _abck, _px3) and send them back from the same identity that earned them; those cookies are bound to the IP, User-Agent, and TLS profile of the session that minted them, and they die when moved.
Two session facts from the field. Deep pagination is where sessions die: on one retailer's category pages, a customer batch ran 99.1% on the US site and 88.6% on the Canadian site because the Canadian crawl went dozens of pages deep on each session and the anti-bot timed the sessions out. And on some Akamai-protected APIs, a 428 with a sec-cp-challenge body is a timed hold, not a block; wait the chlg_duration it names (about 30 seconds) and retry on the same session, and it serves normally. A new session restarts the clock.
Works until: the behavioral engine notices you carry the right cookies but do not behave like a person. Session hygiene defeats stateless detection, not Precursor-style continuous analysis.
8. Avoid honeypots
Some sites plant traps: links hidden with display:none, zero-size elements, or nofollow links no human clicks. Follow one and you have labeled yourself. When you extract links to crawl, filter out elements a rendered browser would treat as invisible, and respect rel="nofollow". In a headless browser, check computed visibility, not just the DOM.
Works until: nothing, really. This is pure downside avoidance. Skipping it gets you thrown out of an easy site.
9. Make every part of the identity agree
The signals are cross-checked, so the failure is usually a contradiction. A US residential IP with Accept-Language: de-DE. A Windows User-Agent from a proxy whose TCP stack looks like Linux. A Chrome User-Agent one major version ahead of the TLS handshake. A timezone from the browser that does not match the proxy's geography. A sec-ch-ua-mobile: ?1 header with a 1920x1080 viewport. Pick one identity per session (OS, browser major, locale, timezone, screen, IP region) and derive every header and every browser setting from it. Fingerprint-generation libraries such as BrowserForge exist for exactly this. Consistency is worth more than novelty.
Works until: the site scores behavior over time, at which point a consistent identity that browses like a script is still a script.
The sequence where this stops working
Here is the pattern teams walk through, in order.
You start with requests and a User-Agent string. It works on a few sites, then you hit your first 403. You add residential proxies. That fixes the reputation blocks and you feel like you solved it. Then you hit Cloudflare and the proxy does nothing, because the problem was your handshake. You switch to curl_cffi with impersonation. More sites open up. Then you hit a site that needs JavaScript, so you move to a stealth browser, and now you maintain a headless fleet and patch it whenever a target learns the current fingerprint. Then a customer says a page you marked green is a challenge page, and you add markers to every target.
Then you hit the wall that does not move: proxies plus a perfect fingerprint plus a stealth browser, and you still land on a challenge, or the page comes back empty and stays empty. Now you are integrating a CAPTCHA solver, tuning per-site behavior, and watching your success rate swing whenever a target updates its bot vendor. The maintenance stops being occasional. It is the job.
This is where the math changes. The open Web Data Frontier Benchmark, run on September 16, 2026 against 100 bot-protected sites (Amazon, Walmart, Zillow, Ticketmaster, and 96 others), shows how wide the gap is even among paid, purpose-built providers: success rates ran from 36.4% to 97.0%, and the median of the 16 APIs was 70.5%. If dedicated vendors whose entire business is unblocking land near 70% on hard targets, a hand-rolled stack will not do better while you also ship your actual product.
Where a managed API fits, honestly
A managed web access API moves the arms race behind someone else's SLA. You send a URL, they handle proxies, fingerprints, rendering, sessions, and challenges, and you get content back. The trade: you give up control and pay per page, in exchange for not maintaining a stealth fleet.
String's Web Access API is the product we make, so treat this paragraph as disclosed bias and check the numbers yourself, because the benchmark is open source and takes your own API keys. It is one HTTP call against any URL with anti-bot handling built in, clean markdown out by default (raw HTML and a JSON envelope also available), and an MCP server so an agent can fetch pages directly. The pricing detail that matters here: String bills per successful request, and a blocked request is not billed (pricing).
curl https://request.usestring.ai/v1/fetch \
-H "Authorization: Bearer $STRING_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "url": "https://example.com/article", "format": "markdown" }'
const r = await fetch("https://request.usestring.ai/v1/fetch", {
method: "POST",
headers: {
Authorization: `Bearer ${process.env.STRING_API_KEY}`,
"Content-Type": "application/json",
},
body: JSON.stringify({ url: "https://hard-target.com", format: "markdown" }),
});
console.log(await r.text());
In the September 16, 2026 run, String passed 97.0% of requests (485 of 500), first of 16 APIs, with the lowest average latency of the 16 at 7.06 seconds (September 2026 results).
Four alternatives worth evaluating on the same run. Scrapfly placed second at 86.2%, credit-based, and its blog is the best public writing on how each anti-bot vendor works. ScraperAPI placed third at 84.0%. Firecrawl passed 80.2%, fourth, with the second-lowest average latency at 9.11 seconds, and it is a genuine crawler: point one call at a domain and it returns every page under it, with markdown output and a self-hostable AGPL core. If the job is crawling cooperative content rather than getting through walls, its benchmark rank is not your binding constraint. Bright Data passed 74.6%, sixth, and brings per-successful-request pricing on Web Unlocker plus enterprise procurement and compliance if you need that.
A decision framework
If your targets are cooperative (docs, blogs, most public content) and your volume is modest, do it yourself. curl_cffi with impersonation, residential proxies, backoff, markers, and caching carries you a long way, and Crawl4AI gives you LLM-ready markdown for free. Keep the money. If you also need JavaScript rendering, add nodriver or Camoufox and cache hard.
If your failures cluster on genuinely hard targets (retail, travel, ticketing, social) and unblocking is a maintenance treadmill, a managed API is cheaper than it looks once you price your engineering time and proxy bandwidth into the comparison. Pick the one that leads the benchmark on the categories you actually hit. At real enterprise scale, with procurement and compliance review, Bright Data is the conventional choice.
The one thing not to do is what most teams do by default: keep patching a homegrown stealth stack past the point where it quietly consumes a full engineer.
FAQ
Why does my scraper get blocked even with rotating proxies?
Because proxies fix one signal. A clean residential IP paired with a Python or Node HTTP client still hands over a non-browser TLS fingerprint, and Cloudflare, DataDome, Akamai, and HUMAN score that fingerprint against your declared User-Agent. Match the TLS and HTTP/2 fingerprint to a real browser (curl_cffi) in addition to using proxies, and rotate per session rather than per request. If you are still blocked after both, the site is using JavaScript or behavioral detection that a stateless client cannot pass.
How do I know if I am blocked or just getting an error?
Do not trust the status code. Anti-bot systems commonly return a 200 with a challenge page, an empty shell, or a real page with zero results. Check the body for a marker that only appears on the real page and treat its absence as a block. Then fetch a page you know works and confirm the marker is present there, so your detector is not lying to you.
Does a stealth browser like Puppeteer with the stealth plugin still work in 2026?
Sometimes, and it is losing ground. The puppeteer-extra-plugin-stealth patches produce a recognizable fingerprint that CreepJS flags. nodriver, zendriver, and Camoufox avoid the WebDriver and patch artifacts that give the plugins away. Even then, Cloudflare treats CDP usage itself as a signal, so no automated browser is fully invisible.
Should I rotate the IP on every request?
Only on stateless targets. On any site that sets cookies, rotate per session and keep one IP per cookie jar. Cloudflare's __cf_bm, DataDome's datadome, and Akamai's _abck cookies are bound to the identity that minted them; changing the IP under them makes you look worse, not better.
Is it legal to scrape a website?
Reading public pages is broadly permitted in the US, but legality depends on where you operate, the data you collect (personal data and copyrighted content carry separate rules) and whether you breach a contract such as a site's terms. This is not legal advice; see is web scraping legal.
How do residential proxies help avoid getting blocked while scraping?
They fix the IP reputation signal. Datacenter ranges are known and scored down; residential and mobile IPs carry real users, so they score neutral. They do not fix the rest: a residential IP with a Python TLS handshake still looks like a bot, and rotating the IP on every request breaks the session score that cookies such as Cloudflare's __cf_bm carry. Rotate per session and match the TLS fingerprint too.
When should I stop building my own scraper and use an API?
When unblocking becomes recurring maintenance rather than a one-time setup: you are patching fingerprints, integrating CAPTCHA solvers, and watching success rates swing every time a target changes vendors. Price your engineering time and proxy bandwidth against per-request API pricing. Purpose-built providers ranged from 36.4% to 97.0% on hard targets in the September 16, 2026 benchmark run; if your stack is well below that and costing you weeks, the trade usually favors the API.
Cheers,
String team
