NewLaunching String Web Access APIRead the manifesto →
← Answers

How to detect a block page that returns HTTP 200

String team · Updated September 30, 2026

Stop trusting the status code and test the body for what the real page must contain. Every response gets three checks: a marker that only the real page carries (a product title, a price field, a listing ID), a count of the items you expected, and a size band learned from known-good fetches. A response that fails any of them is a block, whatever its status. Then match known challenge signatures, so you can tell a block from a layout change.

A 200 only tells you the server sent something back. Login walls, cookie-consent screens, "enable JavaScript" shells and bot challenges all arrive as 200s, and they flow into extraction as if they were data.

What a soft block looks like

We fetched Reddit subreddit pages on September 28, 2026, with a plain Python requests call and with a commercial scraping API (our own). Every response was HTTP 200. Most carried no posts.

Client Page Status Bytes Posts in body
Plain requests r/webscraping 200 8,411 0
Plain requests r/fishing/new 200 8,411 0
Scraping API r/fishing/new 200 361,794 0
Scraping API r/webscraping/new 200 566,641 3

Reddit's own feed fragment for r/fishing returned 22 posts in the same session, so a full page holds roughly 20 or more. The plain client got an 8 KB JavaScript challenge. The API got a 362 KB site shell with an empty feed. Any client, a paid scraping API included, can receive a 200 with no content. A size check passes the 362 KB page. A post count fails it.

The asker in the most-cited Reddit thread on this question lists what most teams try first: password fields, short extracted text, phrases like "sign in" or "accept cookies", and text-to-boilerplate ratios. Each one gives false positives, because an article about authentication says "log in" repeatedly. Negative checks look for signs of a block. Positive checks confirm the content you came for is present, and they are harder to fool.

A detection checklist

  1. Positive marker. Pick a string or selector that exists only on the real page: the product title, a "sku" field in embedded JSON, the listing address. Make it specific to the URL. A generic word can also appear in the challenge page's own markup (see the benchmark section below).
  2. Item count. For lists and feeds, count the items (<article>, result cards, JSON array length) and set a floor. Reddit's 3-post page passes a marker check and fails a floor of 15.
  3. Size band. Record the byte size of known-good responses per page type and flag anything far outside it. Treat size as a supporting check only. The 8 KB challenge is easy to spot, but the 362 KB empty shell falls inside a naive band.
  4. Challenge signatures. Match the vendor's own tells. Cloudflare documents that every challenge page carries the header cf-mitigated: challenge, and its interstitial title reads "Just a moment...". Reddit's pages carried js_challenge or "blocked by network security" in our runs.
  5. Login and consent redirects. Compare the final URL with the one you asked for. A redirect chain ending on /login, /signin, /consent or /account is a wall. A page with a password input and no marker is a login wall. A page with a login prompt and the marker is a normal page.
  6. Structured data. Where the site embeds JSON-LD or a hydration blob, parse it and check required fields. A missing price in a product schema is a clearer failure than a missing CSS class.

Stop at the first failure and record which check failed. That label tells you whether to retry or to fix your parser.

What to do when a check fails

Classify each fetch as ok, blocked (challenge signature or wall), empty (no signature, but the marker or count fails) or error (timeout, 5xx). Retry blocked on a different route: a new IP, a real browser, or a pause. Retry empty once, then send it for review, because it is often a site redesign rather than a block. Cap retries per URL, never write a failed body to your store, and alert on the block rate per domain, since a rising rate is the earliest sign a site changed its protection.

Tools get this wrong in public. A crawl4ai issue from July 2026 reports the Docker API turning a detected anti-bot block into a generic HTTP 500, so the caller cannot tell a blocked site from a broken service. The reporter's summary: "A blocked page is a crawl outcome, not a server fault." A Firecrawl issue from August 2026 describes a job on a protected page that "hangs until the job timeout" and never reaches success or failure. Whatever sits between you and the site, give blocks their own outcome in your code.

How our benchmark scores a success

The Web Data Frontier Benchmark counts a request as a success when the status is 2xx and the body contains a page-specific string, case-insensitive. That logic is validateResponse in src/check.ts of the public harness. It has no block detection of its own. The check is necessary but not sufficient, and our own suite shows why. The Reddit target's marker is webscraping, and Reddit's 8 KB challenge page contains /r/webscraping/ in its form action, so a challenge page passes. Of the 100 markers in the suite, 39 are 15 characters or fewer, such as "NASA" or "London". Short markers are easy to write and easy to match by accident.

If you copy the method, use a long, URL-specific marker with an item count and a signature check beside it. Once a page passes, the record and run checks in how to check that scraped data is complete and accurate take over. The benchmark problem write-up covers why pass criteria decide what a success rate means.

Akamai-protected sites are a common source of 200 challenges; see how to get past Akamai Bot Manager. The best web scraping APIs hub compares providers on the same 100 targets.

FAQ

Why does my scraper get HTTP 200 but no data?

The site served a challenge page, a login or consent wall, or an empty JavaScript shell with a 200 status. The status code only says the server responded. Check the body for a marker that only the real page contains, and count the items you expected.

How do I tell a login wall from a normal page with a sign-in link?

Test for the content you came for first. A page with a password input and no content marker is a wall. A page that has the marker is a real page, even if it also shows a sign-in prompt. Also check whether the final URL redirected to a login or consent path.

Is response size a reliable block signal?

Only as a supporting check. In our Reddit test a plain client got an 8,411-byte challenge page, which a size check catches. A scraping API got a 361,794-byte page with zero posts, which a size check misses. Pair size with a marker and an item count.

How do I detect a Cloudflare challenge in code?

Check the response headers for cf-mitigated: challenge. Cloudflare documents that every challenge page type carries it. The interstitial page title "Just a moment..." is a secondary signal.

Should I retry when I detect a block?

Yes, on a different route: a new IP, a real browser, or after a pause, with a cap per URL. Retry an empty page with no block signature once, then review it, because it often means the site changed its layout.

How does the Web Data Frontier Benchmark decide a request succeeded?

It requires a 2xx status and a page-specific string in the body. The check is public in src/check.ts. It has no separate block detection, so a short marker can match a challenge page, as Reddit's did.

Sources

Get your API key →Explore the Web Access API
© 2026 StringEU and UK GDPR Article 27 representative — appointment verified by EuverifyBuilt in New York City 🗽 🍎