Stop trusting the status code and test the body for what the real page must contain. Every response gets three checks: a marker that only the real page carries (a product title, a price field, a listing ID), a count of the items you expected, and a size band learned from known-good fetches. A response that fails any of them is a block, whatever its status. Then match known challenge signatures, so you can tell a block from a layout change.
A 200 only tells you the server sent something back. Login walls, cookie-consent screens, "enable JavaScript" shells and bot challenges all arrive as 200s, and they flow into extraction as if they were data.
We fetched Reddit subreddit pages on September 28, 2026, with a plain Python requests call and with a commercial scraping API (our own). Every response was HTTP 200. Most carried no posts.
| Client | Page | Status | Bytes | Posts in body |
|---|---|---|---|---|
Plain requests |
r/webscraping | 200 | 8,411 | 0 |
Plain requests |
r/fishing/new | 200 | 8,411 | 0 |
| Scraping API | r/fishing/new | 200 | 361,794 | 0 |
| Scraping API | r/webscraping/new | 200 | 566,641 | 3 |
Reddit's own feed fragment for r/fishing returned 22 posts in the same session, so a full page holds roughly 20 or more. The plain client got an 8 KB JavaScript challenge. The API got a 362 KB site shell with an empty feed. Any client, a paid scraping API included, can receive a 200 with no content. A size check passes the 362 KB page. A post count fails it.
The asker in the most-cited Reddit thread on this question lists what most teams try first: password fields, short extracted text, phrases like "sign in" or "accept cookies", and text-to-boilerplate ratios. Each one gives false positives, because an article about authentication says "log in" repeatedly. Negative checks look for signs of a block. Positive checks confirm the content you came for is present, and they are harder to fool.
"sku" field in embedded JSON, the listing address. Make it specific to the URL. A generic word can also appear in the challenge page's own markup (see the benchmark section below).<article>, result cards, JSON array length) and set a floor. Reddit's 3-post page passes a marker check and fails a floor of 15.cf-mitigated: challenge, and its interstitial title reads "Just a moment...". Reddit's pages carried js_challenge or "blocked by network security" in our runs./login, /signin, /consent or /account is a wall. A page with a password input and no marker is a login wall. A page with a login prompt and the marker is a normal page.Stop at the first failure and record which check failed. That label tells you whether to retry or to fix your parser.
Classify each fetch as ok, blocked (challenge signature or wall), empty (no signature, but the marker or count fails) or error (timeout, 5xx). Retry blocked on a different route: a new IP, a real browser, or a pause. Retry empty once, then send it for review, because it is often a site redesign rather than a block. Cap retries per URL, never write a failed body to your store, and alert on the block rate per domain, since a rising rate is the earliest sign a site changed its protection.
Tools get this wrong in public. A crawl4ai issue from July 2026 reports the Docker API turning a detected anti-bot block into a generic HTTP 500, so the caller cannot tell a blocked site from a broken service. The reporter's summary: "A blocked page is a crawl outcome, not a server fault." A Firecrawl issue from August 2026 describes a job on a protected page that "hangs until the job timeout" and never reaches success or failure. Whatever sits between you and the site, give blocks their own outcome in your code.
The Web Data Frontier Benchmark counts a request as a success when the status is 2xx and the body contains a page-specific string, case-insensitive. That logic is validateResponse in src/check.ts of the public harness. It has no block detection of its own. The check is necessary but not sufficient, and our own suite shows why. The Reddit target's marker is webscraping, and Reddit's 8 KB challenge page contains /r/webscraping/ in its form action, so a challenge page passes. Of the 100 markers in the suite, 39 are 15 characters or fewer, such as "NASA" or "London". Short markers are easy to write and easy to match by accident.
If you copy the method, use a long, URL-specific marker with an item count and a signature check beside it. Once a page passes, the record and run checks in how to check that scraped data is complete and accurate take over. The benchmark problem write-up covers why pass criteria decide what a success rate means.
Akamai-protected sites are a common source of 200 challenges; see how to get past Akamai Bot Manager. The best web scraping APIs hub compares providers on the same 100 targets.
The site served a challenge page, a login or consent wall, or an empty JavaScript shell with a 200 status. The status code only says the server responded. Check the body for a marker that only the real page contains, and count the items you expected.
Test for the content you came for first. A page with a password input and no content marker is a wall. A page that has the marker is a real page, even if it also shows a sign-in prompt. Also check whether the final URL redirected to a login or consent path.
Only as a supporting check. In our Reddit test a plain client got an 8,411-byte challenge page, which a size check catches. A scraping API got a 361,794-byte page with zero posts, which a size check misses. Pair size with a marker and an item count.
Check the response headers for cf-mitigated: challenge. Cloudflare documents that every challenge page type carries it. The interstitial page title "Just a moment..." is a secondary signal.
Yes, on a different route: a new IP, a real browser, or after a pause, with a cap per URL. Retry an empty page with no block signature once, then review it, because it often means the site changed its layout.
It requires a 2xx status and a page-specific string in the body. The check is public in src/check.ts. It has no separate block detection, so a short marker can match a challenge page, as Reddit's did.
requests and a scraping API against Reddit subreddit pages (status, bytes and <shreddit-post> counts as reported above).src/check.ts (validateResponse) and src/tests.const.ts (100 targets and markers).cf-mitigated header).