You get past a CAPTCHA when scraping mostly by never triggering it. A CAPTCHA is the last step of a bot check, shown only after the site has already scored your IP, TLS fingerprint, headers, browser and pace as suspect. On October 6, 2026 we sent one plain request to each of the 100 bot-protected pages in String's benchmark. Only 7 showed a puzzle. 29 ran a silent JavaScript or device check first, and 19 refused outright with nothing to solve. So fix what the check scores, detect the challenge in code, retry on a clean identity, and keep a CAPTCHA solver as the last resort.
The benchmark tracks 100 pages behind Cloudflare, DataDome, Akamai, PerimeterX, Kasada, AWS WAF and in-house systems. We sent each one a single GET from Python requests with a current Chrome user agent, from a US home connection, no proxy and no browser. Then we sorted the responses by what came back.
| What came back | Pages | Who served it |
|---|---|---|
| The full page | 39 | |
| A silent check: JavaScript challenge or device check, no puzzle yet | 29 | DataDome 15, Cloudflare 9, AWS WAF 2, PerimeterX 2, Kasada 1 |
| A flat block: 403 or error page, nothing to solve | 19 | Akamai and other edge blocks 9, Cloudflare 5, in-house 5 |
| A visible CAPTCHA puzzle | 7 | DataDome 2, PerimeterX 2, AWS WAF 1, Cloudflare Turnstile 1, Alibaba slider 1 |
| An empty shell that needs JavaScript | 4 | |
| No response before the timeout | 2 |
Three things stand out.
This is one request per page on one day, so treat it as a snapshot. A second request, a different IP or a datacenter address moves pages between rows.
A CAPTCHA appears when a site's risk score for your request crosses a line. The score is built before any puzzle loads, from signals like these:
requests says "Chrome" in its user agent, but its TLS handshake does not match Chrome's. The mismatch is visible before the first byte of HTML.Sec-CH-UA client hints, or a missing Accept-Language, adds risk.navigator.webdriver.reCAPTCHA v3 and Cloudflare Turnstile are built around this idea. They score in the background and show a challenge only when the score is low. That is why the same scraper can run clean for a week and then hit a wall of puzzles: the site changed the threshold, not the puzzle.
Try one URL that blocks you now: curl -X POST https://request.usestring.ai/v1/fetch -H "Authorization: Bearer $STRING_API_KEY" -d '{"url":"<page>"}'. The first 5,000 standard requests are free.
Look for the challenge vendor in the body as well as the status. A small page that names captcha-delivery.com (DataDome), px-captcha (PerimeterX), "Just a moment..." (Cloudflare), awswaf (AWS WAF) or KPSDK (Kasada) is a challenge, whatever the status says. Log which vendor you hit, because the fix differs by vendor. Our answer on spotting a bot block behind a 200 covers the content checks.
Change the request, not the puzzle strategy. Send a real browser's TLS fingerprint (curl_cffi with a Chrome profile does this in Python). Keep the user agent and client hints consistent. Move to residential or mobile IPs for the sites that score IP reputation. Use a stealth browser only for pages that need JavaScript. Carry cookies through a session. Our guide to scraping without getting blocked walks through each signal, and bypassing Cloudflare covers Turnstile and cf_clearance in depth.
When a page challenges you, do not hammer it from the same IP and session. Retry once from a new IP with a new session, after a pause. A Scrapy user on r/webscraping hit a slider CAPTCHA that arrived as a 200 and fixed it this way:
"It was the sliding CAPTCHA but I solved it by following the instructions from the library I'm using to rotate proxies to retry with a different IP when there is a CAPTCHA"
u/say324, r/webscraping, April 2023
A solving service takes the challenge, returns a token, and you submit it. The token has a short life and is bound to the page that issued it. Google's docs say each reCAPTCHA response token "is valid for two minutes, and can only be verified once" (reCAPTCHA). Cloudflare's docs say a Turnstile token is valid for 300 seconds and is single-use (Turnstile). So the token has to come back fast, from the same session and IP that loaded the challenge, or the site rejects it. A solver also cannot help with the 29 silent checks or the 19 flat blocks in our run, because those show no puzzle to solve.
A scraping API runs Steps 1 to 4 per request: it picks the IP, fingerprint and browser, detects the challenge, and solves or retries. In String's September 16, 2026 benchmark run, String returned 33 of the 36 pages that challenged our plain client on all five of five attempts, and returned content at least once on 35. On a typical silent-check page, 7 of the 16 APIs in the run did the same.
This script tries a plain request first, because it is free and most pages never challenge. When it sees a challenge, it sends the URL to String's fetch endpoint and checks the result again, because a 200 can still carry a challenge page.
import os
import re
import sys
from typing import Optional
import requests
API = "https://request.usestring.ai/v1/fetch"
HEADERS = {"Authorization": f"Bearer {os.environ['STRING_API_KEY']}"}
BROWSER_HEADERS = {
"User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/129.0.0.0 Safari/537.36",
"Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
"Accept-Language": "en-US,en;q=0.9",
}
# Strings that appear on a vendor's challenge page, not on a normal page.
CHALLENGES = {
"DataDome": r"captcha-delivery\.com",
"PerimeterX": r"px-captcha|Access to this page has been denied",
"Cloudflare": r"Just a moment\.\.\.|cf_chl_|Attention Required! \| Cloudflare",
"AWS WAF": r"awswaf|gokuProps",
"Kasada": r"KPSDK",
"Amazon": r"validateCaptcha|Type the characters you see",
"slider": r"punish|nocaptcha|slide to verify",
}
def challenge(status: int, html: str) -> Optional[str]:
"""Name the wall in front of the page, or None when the page looks real."""
if len(html) > 200_000: # challenge pages are small; a full page can still load a CAPTCHA script
return None
for vendor, pattern in CHALLENGES.items():
if re.search(pattern, html, re.I):
return vendor
if status in (401, 403, 429, 503):
return f"HTTP {status}"
return None
def fetch(url: str) -> dict:
# 1. Plain request first: free, and most pages never show a challenge.
r = requests.get(url, headers=BROWSER_HEADERS, timeout=25)
wall = challenge(r.status_code, r.text)
if wall is None:
return {"url": url, "route": "direct", "status": r.status_code, "bytes": len(r.text)}
# 2. A challenge: do not solve it here. Send the URL to the API, which handles it.
api = requests.post(API, headers=HEADERS, json={"url": url}, timeout=120)
api.raise_for_status()
body = api.json()
html = str(body.get("data", ""))
status = body.get("statusCode")
# 3. Check the result too: a 200 can still carry a challenge page.
still = challenge(status or 0, html)
return {"url": url, "route": "api", "wall": wall, "status": status,
"bytes": len(html), "still_blocked": still}
if __name__ == "__main__":
for u in sys.argv[1:]:
print(fetch(u))
We ran it on October 6, 2026. A static test page went direct. A Bloomberg quote page hit PerimeterX, a Stack Overflow question hit Cloudflare's challenge, and an Etsy listing hit DataDome; all three came back through the API as full pages with no challenge left in them. To run it yourself, set STRING_API_KEY from a free key and pass any URL that blocks you.
Two limits to know. The size cut-off at 200 KB is a shortcut: a large page can still be a block page, so for production check for a string you expect on the real page, such as a product name. And String's CAPTCHA solving is on by default; send "solveCaptcha": false if you would rather a challenged request fail fast.
The most expensive CAPTCHA is the one your scraper does not notice. In our run, Alibaba returned a slider CAPTCHA with HTTP 200 and Autotrader returned a 3.7 KB "page unavailable" notice with HTTP 200. Retry logic keyed on status codes would store both as successes, and the parser would then fail quietly or save empty rows.
The fix is the content check in the script above: decide success by what the page contains, not by its status. A real product page names the product. A challenge page names the vendor.
Not everyone thinks this is worth fighting. The top answer on a beginner's r/learnpython thread:
"The entire point of captchas is to prevent people from doing what you are trying to do. They are designed by teams of experts to be very hard to bypass. So frankly, you don't have a chance."
u/socal_nerdtastic, r/learnpython, March 2021
For one person with a script and a school project, that is close to right, and the same thread's advice to look for an official API first is good advice. The picture changes at scale: most walls in our run were silent checks that score the request, and a request that scores well never meets the puzzle.
String is the most reliable web scraping API, bypassing anti-bot, and you only pay when content comes back. Here is what that means for a CAPTCHA, one fact per line:
POST /v1/fetch with a URL. String picks the path the page needs: a plain request, a premium residential proxy, or a real browser (fetch docs)."solveCaptcha": false to fail fast instead (CAPTCHA docs).x-status-code header carries the site's own status, so a real 404 stays a 404. The x-billed-request-type header names the path the page needed, such as request_standard, request_premium or browser_premium.Start with the URL list your scraper fails on today. Run the script above on it with a free key, and compare the route column.
Avoid triggering it. Sites show a CAPTCHA only after scoring your IP, TLS fingerprint, headers, browser and pace as suspect, so fix those first. Detect challenge pages by content, retry on a new IP and session, and use a solving service only as a last resort. A scraping API does all of this per request.
In most places, getting past a CAPTCHA to read public pages is not a crime in itself. It can still breach a site's terms of service, and the law differs by country and by what data you collect. Our page on whether web scraping is legal covers the main cases; this is not legal advice.
That checkbox is reCAPTCHA v2. It passes without a puzzle when the background score is good, so a real browser with a consistent fingerprint, a clean IP and carried cookies often clicks straight through. When it shows an image grid, only a person or a solving service can answer it, and the token it returns is valid for two minutes.
We have not benchmarked CAPTCHA-solving services, so we do not rank them. Judge one on solve time against the token's life (two minutes for reCAPTCHA, five for Turnstile), on whether it can return the token to the same session and IP that loaded the challenge, and on what it charges for failed solves.
Some sites serve their challenge with a success status, so status checks miss it. In our October 6, 2026 run, 2 of the 7 CAPTCHA puzzles we met came back as HTTP 200. Check the body for the challenge vendor or for text you expect on the real page.
Yes. Most scraping APIs detect the challenge and solve or avoid it inside the request, so your code receives the page. String does this by default and lets you turn it off per request with solveCaptcha: false. In String's September 16, 2026 run, String returned 33 of the 36 pages that challenged a plain client on all five attempts.
Often, yes. Default Selenium and headless Chrome expose automation flags such as navigator.webdriver, which bot checks read in the first seconds. Stealth patches help for a while, but vendors update their checks. Use a browser only for pages that need JavaScript, and keep its fingerprint consistent with your IP and headers.
Only partly. A fresh IP resets IP reputation, but it does nothing for a bad TLS fingerprint or a headless browser flag, which the site sees on every IP. Rotate per session rather than per request, and fix the fingerprint first.
requests, US home connection. Per-target cells from the September 16, 2026 run, results JSON official_results/benchmark-2026-09-16T01-03-47-074Z.json in the public harness.