NewLaunching String Web Access APIRead the manifesto →
← Answers

How to get past CAPTCHA when scraping

String team · Updated October 6, 2026

You get past a CAPTCHA when scraping mostly by never triggering it. A CAPTCHA is the last step of a bot check, shown only after the site has already scored your IP, TLS fingerprint, headers, browser and pace as suspect. On October 6, 2026 we sent one plain request to each of the 100 bot-protected pages in String's benchmark. Only 7 showed a puzzle. 29 ran a silent JavaScript or device check first, and 19 refused outright with nothing to solve. So fix what the check scores, detect the challenge in code, retry on a clean identity, and keep a CAPTCHA solver as the last resort.

What 100 bot-protected pages showed a plain client

The benchmark tracks 100 pages behind Cloudflare, DataDome, Akamai, PerimeterX, Kasada, AWS WAF and in-house systems. We sent each one a single GET from Python requests with a current Chrome user agent, from a US home connection, no proxy and no browser. Then we sorted the responses by what came back.

What came back Pages Who served it
The full page 39
A silent check: JavaScript challenge or device check, no puzzle yet 29 DataDome 15, Cloudflare 9, AWS WAF 2, PerimeterX 2, Kasada 1
A flat block: 403 or error page, nothing to solve 19 Akamai and other edge blocks 9, Cloudflare 5, in-house 5
A visible CAPTCHA puzzle 7 DataDome 2, PerimeterX 2, AWS WAF 1, Cloudflare Turnstile 1, Alibaba slider 1
An empty shell that needs JavaScript 4
No response before the timeout 2

Three things stand out.

  • The puzzle is the rare case. 7 of 100 pages showed one. Four times as many ran a check you never see.
  • The CAPTCHA is armed on pages that let you in. 11 of the 39 pages that served full content still loaded a CAPTCHA or challenge script, such as reCAPTCHA, hCaptcha, GeeTest, Turnstile, AWS WAF or Kasada. The site held it back because the request scored well enough.
  • Two of the 7 puzzles came back as HTTP 200. Alibaba served a slider CAPTCHA and Mouser served a PerimeterX denial page with a CAPTCHA, both with a success status. A scraper that checks only the status code stores those as data.

This is one request per page on one day, so treat it as a snapshot. A second request, a different IP or a datacenter address moves pages between rows.

Why the CAPTCHA appears

A CAPTCHA appears when a site's risk score for your request crosses a line. The score is built before any puzzle loads, from signals like these:

  • IP reputation. Datacenter ranges score worse than home and mobile addresses. Our run came from a home connection and still drew 55 checks, blocks or puzzles.
  • TLS and HTTP/2 fingerprint. Python requests says "Chrome" in its user agent, but its TLS handshake does not match Chrome's. The mismatch is visible before the first byte of HTML.
  • Headers. A user agent that disagrees with the Sec-CH-UA client hints, or a missing Accept-Language, adds risk.
  • Browser. Default headless Chrome, Selenium and Puppeteer expose automation flags such as navigator.webdriver.
  • Pace and session. Bursts, a fresh session on every request, and no cookies carried between pages all look automated.

reCAPTCHA v3 and Cloudflare Turnstile are built around this idea. They score in the background and show a challenge only when the score is low. That is why the same scraper can run clean for a week and then hit a wall of puzzles: the site changed the threshold, not the puzzle.

Try one URL that blocks you now: curl -X POST https://request.usestring.ai/v1/fetch -H "Authorization: Bearer $STRING_API_KEY" -d '{"url":"<page>"}'. The first 5,000 standard requests are free.

How to get past it, step by step

Step 1: Detect the wall, not just the status code

Look for the challenge vendor in the body as well as the status. A small page that names captcha-delivery.com (DataDome), px-captcha (PerimeterX), "Just a moment..." (Cloudflare), awswaf (AWS WAF) or KPSDK (Kasada) is a challenge, whatever the status says. Log which vendor you hit, because the fix differs by vendor. Our answer on spotting a bot block behind a 200 covers the content checks.

Step 2: Fix what the check scores

Change the request, not the puzzle strategy. Send a real browser's TLS fingerprint (curl_cffi with a Chrome profile does this in Python). Keep the user agent and client hints consistent. Move to residential or mobile IPs for the sites that score IP reputation. Use a stealth browser only for pages that need JavaScript. Carry cookies through a session. Our guide to scraping without getting blocked walks through each signal, and bypassing Cloudflare covers Turnstile and cf_clearance in depth.

Step 3: Retry on a clean identity

When a page challenges you, do not hammer it from the same IP and session. Retry once from a new IP with a new session, after a pause. A Scrapy user on r/webscraping hit a slider CAPTCHA that arrived as a 200 and fixed it this way:

"It was the sliding CAPTCHA but I solved it by following the instructions from the library I'm using to rotate proxies to retry with a different IP when there is a CAPTCHA"

u/say324, r/webscraping, April 2023

Step 4: Use a solver only as the last resort

A solving service takes the challenge, returns a token, and you submit it. The token has a short life and is bound to the page that issued it. Google's docs say each reCAPTCHA response token "is valid for two minutes, and can only be verified once" (reCAPTCHA). Cloudflare's docs say a Turnstile token is valid for 300 seconds and is single-use (Turnstile). So the token has to come back fast, from the same session and IP that loaded the challenge, or the site rejects it. A solver also cannot help with the 29 silent checks or the 19 flat blocks in our run, because those show no puzzle to solve.

Step 5: Or hand the URL to a scraping API

A scraping API runs Steps 1 to 4 per request: it picks the IP, fingerprint and browser, detects the challenge, and solves or retries. In String's September 16, 2026 benchmark run, String returned 33 of the 36 pages that challenged our plain client on all five of five attempts, and returned content at least once on 35. On a typical silent-check page, 7 of the 16 APIs in the run did the same.

Python: detect the wall and fall back

This script tries a plain request first, because it is free and most pages never challenge. When it sees a challenge, it sends the URL to String's fetch endpoint and checks the result again, because a 200 can still carry a challenge page.

import os
import re
import sys
from typing import Optional

import requests

API = "https://request.usestring.ai/v1/fetch"
HEADERS = {"Authorization": f"Bearer {os.environ['STRING_API_KEY']}"}
BROWSER_HEADERS = {
    "User-Agent": "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 "
                  "(KHTML, like Gecko) Chrome/129.0.0.0 Safari/537.36",
    "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
    "Accept-Language": "en-US,en;q=0.9",
}

# Strings that appear on a vendor's challenge page, not on a normal page.
CHALLENGES = {
    "DataDome": r"captcha-delivery\.com",
    "PerimeterX": r"px-captcha|Access to this page has been denied",
    "Cloudflare": r"Just a moment\.\.\.|cf_chl_|Attention Required! \| Cloudflare",
    "AWS WAF": r"awswaf|gokuProps",
    "Kasada": r"KPSDK",
    "Amazon": r"validateCaptcha|Type the characters you see",
    "slider": r"punish|nocaptcha|slide to verify",
}


def challenge(status: int, html: str) -> Optional[str]:
    """Name the wall in front of the page, or None when the page looks real."""
    if len(html) > 200_000:  # challenge pages are small; a full page can still load a CAPTCHA script
        return None
    for vendor, pattern in CHALLENGES.items():
        if re.search(pattern, html, re.I):
            return vendor
    if status in (401, 403, 429, 503):
        return f"HTTP {status}"
    return None


def fetch(url: str) -> dict:
    # 1. Plain request first: free, and most pages never show a challenge.
    r = requests.get(url, headers=BROWSER_HEADERS, timeout=25)
    wall = challenge(r.status_code, r.text)
    if wall is None:
        return {"url": url, "route": "direct", "status": r.status_code, "bytes": len(r.text)}

    # 2. A challenge: do not solve it here. Send the URL to the API, which handles it.
    api = requests.post(API, headers=HEADERS, json={"url": url}, timeout=120)
    api.raise_for_status()
    body = api.json()
    html = str(body.get("data", ""))
    status = body.get("statusCode")
    # 3. Check the result too: a 200 can still carry a challenge page.
    still = challenge(status or 0, html)
    return {"url": url, "route": "api", "wall": wall, "status": status,
            "bytes": len(html), "still_blocked": still}


if __name__ == "__main__":
    for u in sys.argv[1:]:
        print(fetch(u))

We ran it on October 6, 2026. A static test page went direct. A Bloomberg quote page hit PerimeterX, a Stack Overflow question hit Cloudflare's challenge, and an Etsy listing hit DataDome; all three came back through the API as full pages with no challenge left in them. To run it yourself, set STRING_API_KEY from a free key and pass any URL that blocks you.

Two limits to know. The size cut-off at 200 KB is a shortcut: a large page can still be a block page, so for production check for a string you expect on the real page, such as a product name. And String's CAPTCHA solving is on by default; send "solveCaptcha": false if you would rather a challenged request fail fast.

The failure mode: a CAPTCHA inside a 200

The most expensive CAPTCHA is the one your scraper does not notice. In our run, Alibaba returned a slider CAPTCHA with HTTP 200 and Autotrader returned a 3.7 KB "page unavailable" notice with HTTP 200. Retry logic keyed on status codes would store both as successes, and the parser would then fail quietly or save empty rows.

The fix is the content check in the script above: decide success by what the page contains, not by its status. A real product page names the product. A challenge page names the vendor.

The honest counterpoint

Not everyone thinks this is worth fighting. The top answer on a beginner's r/learnpython thread:

"The entire point of captchas is to prevent people from doing what you are trying to do. They are designed by teams of experts to be very hard to bypass. So frankly, you don't have a chance."

u/socal_nerdtastic, r/learnpython, March 2021

For one person with a script and a school project, that is close to right, and the same thread's advice to look for an official API first is good advice. The picture changes at scale: most walls in our run were silent checks that score the request, and a request that scores well never meets the puzzle.

How String gets past CAPTCHAs

String is the most reliable web scraping API, bypassing anti-bot, and you only pay when content comes back. Here is what that means for a CAPTCHA, one fact per line:

  • One call per URL. You send POST /v1/fetch with a URL. String picks the path the page needs: a plain request, a premium residential proxy, or a real browser (fetch docs).
  • CAPTCHAs are solved inside the request. CAPTCHA solving is on by default, so your code gets the page, not the puzzle. Send "solveCaptcha": false to fail fast instead (CAPTCHA docs).
  • The response says what happened. The x-status-code header carries the site's own status, so a real 404 stays a 404. The x-billed-request-type header names the path the page needed, such as request_standard, request_premium or browser_premium.
  • A blocked request costs nothing. String bills only when content comes back. A CAPTCHA that String cannot pass is not charged.
  • The result is measured. In the September 16, 2026 benchmark run, String returned content on 97.0% of 500 requests to 100 bot-protected pages, the highest of 16 APIs. It returned 33 of the 36 pages that challenged our plain client on all five attempts.
  • Agents get the same fetch as a tool. String's hosted MCP server gives Claude and other MCP clients search and fetch, with CAPTCHA handling included.

Where String fits

  • For: teams that need public pages from sites behind DataDome, Cloudflare, PerimeterX, Akamai or Kasada, and would rather not run proxies, fingerprints and solvers themselves.
  • Not for: pages behind a login you do not own, or a site that offers the same data through an official API.
  • Price: you pay only when content comes back. On the Starter plan ($20 a month), a plain fetch is $0.30 per 1,000 on standard proxies and $3.00 on premium; a browser fetch is $1.50 and $6.00. The pricing page lists no separate CAPTCHA charge. The first 5,000 standard requests are free.
  • Integrations: a REST API you can call from any language, and an MCP server for agents.
  • Limits: 60 requests a second per account with a 3,600-request burst, no cap on concurrency.

Start with the URL list your scraper fails on today. Run the script above on it with a free key, and compare the route column.

FAQ

How to get past CAPTCHA when scraping?

Avoid triggering it. Sites show a CAPTCHA only after scoring your IP, TLS fingerprint, headers, browser and pace as suspect, so fix those first. Detect challenge pages by content, retry on a new IP and session, and use a solving service only as a last resort. A scraping API does all of this per request.

Is bypassing CAPTCHA illegal?

In most places, getting past a CAPTCHA to read public pages is not a crime in itself. It can still breach a site's terms of service, and the law differs by country and by what data you collect. Our page on whether web scraping is legal covers the main cases; this is not legal advice.

How do I bypass the "I am not a robot" CAPTCHA?

That checkbox is reCAPTCHA v2. It passes without a puzzle when the background score is good, so a real browser with a consistent fingerprint, a clean IP and carried cookies often clicks straight through. When it shows an image grid, only a person or a solving service can answer it, and the token it returns is valid for two minutes.

What is the best CAPTCHA-solving service for web scraping?

We have not benchmarked CAPTCHA-solving services, so we do not rank them. Judge one on solve time against the token's life (two minutes for reCAPTCHA, five for Turnstile), on whether it can return the token to the same session and IP that loaded the challenge, and on what it charges for failed solves.

Why does my scraper get a CAPTCHA page with a 200 status code?

Some sites serve their challenge with a success status, so status checks miss it. In our October 6, 2026 run, 2 of the 7 CAPTCHA puzzles we met came back as HTTP 200. Check the body for the challenge vendor or for text you expect on the real page.

Can a scraping API solve CAPTCHAs automatically?

Yes. Most scraping APIs detect the challenge and solve or avoid it inside the request, so your code receives the page. String does this by default and lets you turn it off per request with solveCaptcha: false. In String's September 16, 2026 run, String returned 33 of the 36 pages that challenged a plain client on all five attempts.

Does Selenium trigger CAPTCHAs?

Often, yes. Default Selenium and headless Chrome expose automation flags such as navigator.webdriver, which bot checks read in the first seconds. Stealth patches help for a while, but vendors update their checks. Use a browser only for pages that need JavaScript, and keep its fingerprint consistent with your IP and headers.

Do rotating proxies stop CAPTCHAs?

Only partly. A fresh IP resets IP reputation, but it does nothing for a bad TLS fingerprint or a headless browser flag, which the site sees on every IP. Rotate per session rather than per request, and fix the fingerprint first.

Sources

Get your API key →Explore the Web Access API
© 2026 StringEU and UK GDPR Article 27 representative — appointment verified by EuverifyBuilt in New York City 🗽 🍎