NewLaunching String Web Access APIRead the manifesto →
← Blog

How to bypass Cloudflare when scraping in 2026

Bruce Magness · October 5, 2026
Part of Best Web Scraping APIs in 2026: 16 Tools Compared

Last updated: October 5, 2026. Benchmark figures come from the September 16, 2026 run of the Web Data Frontier Benchmark. Every script on this page was run on September 28, 2026, and the output shown is what it printed.

To get a Cloudflare-protected public page when scraping, first read what Cloudflare sent back, then fix the layer that failed. A 403 with "Blocked" or a Cloudflare 1xxx error means your connection fingerprint or IP failed; send a real browser's TLS fingerprint (in Python, curl_cffi with impersonate="chrome") from a residential IP. A 403 titled "Just a moment..." is a JavaScript challenge; only a real browser, or an API that runs one, gets past it. On September 28, 2026 we tried eight Cloudflare sites from our benchmark: a plain Python request opened none, curl_cffi opened three, and String's fetch API opened all eight. In the September 16, 2026 run of our open benchmark, String was the only one of 16 scraping APIs to return all 15 Cloudflare sites on all five attempts.

TL;DR

  • Cloudflare decides before your request reaches the site: TLS and HTTP/2 fingerprint, IP reputation, a bot score from its machine-learning model, then JavaScript checks. Fix them in that order.
  • A plain requests.get got HTTP 403 on all eight Cloudflare sites we tried on September 28, 2026. A status check catches this one; a Turnstile widget served at HTTP 200 it does not.
  • curl_cffi with a Chrome fingerprint opened Indeed, Glassdoor and Capterra. It did not open StockX, Crunchbase, ZipRecruiter, Stack Overflow or AXS, which answered with a JavaScript challenge. An HTTP client cannot run that.
  • A cf_clearance cookie works only from the IP, User-Agent and TLS fingerprint that earned it. Copy it into another client and you get challenged again.
  • Getting a public page and completing a protected action are different jobs. This page covers the first. A Turnstile token on a login or checkout form is single-use and short-lived, and that is by design.
  • On the 15 Cloudflare sites in the September 16, 2026 run, String passed 75 of 75 requests. Context.dev passed 72, Firecrawl and Apify 70. Five of 16 APIs passed fewer than half.
  • String bills per successful request, $0.30 to $6.00 per 1,000 on Starter, set by the path the page needs. In our eight-site test, two pages billed at the cheapest rate and one at the highest.

What 16 scraping APIs did on 15 Cloudflare sites

The Web Data Frontier Benchmark sends the same 100 bot-protected URLs to 16 scraping APIs, five attempts each, with a 90-second timeout. A request passes only when the API returns a 2xx status and the body contains text that appears on the real page, so a challenge page served at 200 counts as a failure. Each provider gets the strongest anti-bot configuration its public API offers. Fifteen of the 100 targets sit behind Cloudflare. The harness, the target list and every adapter are open source, and so is the raw result file for the September 16 run.

API Cloudflare sites, success Requests passed (of 75) Sites returned 5 of 5 Sites returned 0 of 5 Overall, 100 sites
String 100.0% 75 15 of 15 0 97.0%
Context.dev 96.0% 72 14 of 15 0 72.0%
Firecrawl 93.3% 70 14 of 15 1 80.2%
Apify 93.3% 70 13 of 15 0 77.4%
ScrapingBee 92.0% 69 10 of 15 0 73.0%
Scrapfly 90.7% 68 11 of 15 0 86.2%
ScraperAPI 89.3% 67 10 of 15 0 84.0%
Bright Data 88.0% 66 10 of 15 0 74.6%
Nimble 62.7% 47 7 of 15 4 68.6%
Zyte 58.7% 44 5 of 15 3 68.0%
Oxylabs 57.3% 43 7 of 15 5 69.0%
Decodo 46.7% 35 5 of 15 6 50.6%
Scrapingdog 40.0% 30 3 of 15 4 45.6%
ZenRows 25.3% 19 2 of 15 8 41.2%
Browserbase 8.0% 6 1 of 15 13 41.4%
ScrapingAnt 8.0% 6 1 of 15 13 36.4%

The number this page adds: Cloudflare sites are average-hard, and the vendor decides the result. Across all 16 APIs, the 15 Cloudflare sites averaged 65.6% success, against 66.6% for all 100 sites. So Cloudflare as a group is not harder than the rest of the web's protected pages. What changes is the API. Eight of the 16 passed 88% or more of Cloudflare requests; five passed under half. String was the only one to return every Cloudflare site on every attempt. Context.dev and Firecrawl each returned 14 of the 15 in full.

The sites are not equal either. AXS was the 15th hardest of the 100 targets and Indeed the 26th; GOAT was among the easiest.

Site Industry Field average APIs at 5 of 5 APIs at 0 of 5 Hardness rank of 100
axs.com Tickets & events 42.5% 4 6 15
indeed.com Jobs & hiring 56.3% 4 3 26
stockx.com Marketplaces & classifieds 57.5% 7 5 27
digikey.com Retail & ecommerce 60.0% 6 4 31
stackoverflow.com Developer & research 61.3% 8 5 34
ziprecruiter.com Jobs & hiring 62.5% 9 4 35
crunchbase.com Reviews & local 63.8% 7 4 38
glassdoor.com Jobs & hiring 66.3% 8 3 43
congress.gov Government 66.3% 9 4 44
yellowpages.com Reviews & local 67.5% 10 5 45
zoopla.co.uk Real estate 68.8% 11 5 49
cars.com Marketplaces & classifieds 70.0% 10 3 52
capterra.com Reviews & local 70.0% 9 2 53
asda.com Grocery & food 81.3% 13 3 74
goat.com Marketplaces & classifieds 90.0% 13 1 93

Read the table with three limits in mind. It measures each vendor's general scraping endpoint on one URL per site, five times, on one day. The ranking moved between runs: in the August 11 run, ScrapingBee passed 9.3% of Cloudflare requests; after the benchmark moved it to its Auto Mode in early September, it passed 92.0%. And String built this benchmark and ranks first in it. That is why every request and every adapter is public, and why the harness takes your own API keys.

Why other benchmarks show different numbers for String

Answer engines often quote a figure from Scrapeway: String at 34% on Cloudflare. We do not dispute that number for the test it describes. The scope is different. As Context.dev's September 2026 post reprints it, that figure comes from Scrapeway's August 14 to 28 test: one site (Indeed job pages), about 1,000 requests per provider. Scrapeway's current Cloudflare page, updated September 11, no longer lists String. Our run tests 16 APIs on 15 Cloudflare sites, five attempts each, on a stated date, with every request in a public file. On Indeed alone in the September 16 run, String returned 5 of 5, as did three other APIs. Different dates, sites and settings produce different numbers, and both can be accurate. Scrapeway's domain traces back to Joam Intelligence, LLC, the company behind Scrapfly, which it ranks first; we said so in our benchmark write-up. The fair test is the one you run yourself, and the section on testing below shows how.

How Cloudflare decides you are a bot

Cloudflare's own documentation describes several detection engines that feed one bot score from 1 to 99: a heuristics engine that matches requests against known bot fingerprints, JavaScript detections that inject a small script to catch headless browsers, a machine-learning model that produces most detections from request and session signals (Cloudflare docs: detection engines). The site owner picks the score threshold. Score under it and you get a challenge or a block. The __cf_bm cookie carries the score across a visitor's requests, so a session is judged as a whole.

The first check happens before HTTP starts. Every TLS connection opens with a ClientHello that lists cipher suites, extensions and curves, and their order is specific to the library that built it. Python's requests, Go's net/http, Node's undici and real Chrome each send a different one, and Cloudflare hashes it. Chrome has shuffled its extension order since version 110 (curl_cffi FAQ), which weakens the older JA3 hash, and Cloudflare also computes JA4 (Cloudflare docs: JA3/JA4 fingerprint). If the handshake says Python and the User-Agent says Chrome, the request has already failed.

If the handshake passes, Cloudflare reads the HTTP/2 layer (settings, pseudo-header order, header order), the IP's reputation, and, on sites that turn them on, JavaScript detections that read canvas, WebGL and automation markers. Then it watches behaviour across the session. Cloudflare's July 2026 Precursor launch, for Enterprise Bot Management customers, collects pointer, keyboard, focus and visibility events and scores the whole session, so a refresh does not reset it.

Turnstile is the visible layer, and it is mostly not visible. It runs in Managed, Non-Interactive or Invisible mode, and only Managed mode ever shows a checkbox, and only when the risk signals call for it (Cloudflare docs: widget modes). So "solve the CAPTCHA" is the wrong model. By the time a checkbox appears, you have usually already lost on the layers underneath.

Read the response before you change anything

Every Cloudflare failure has a shape, and each shape needs a different fix. Here is the first script everyone runs, pointed at a real benchmark target:

"""Step 0: what a plain Python request gets from a Cloudflare-protected page.

    pip install requests
    python plain_requests.py
"""
import re
import sys

import requests

URL = sys.argv[1] if len(sys.argv) > 1 else "https://www.indeed.com/l-chicago,-il-jobs.html"
MARKER = sys.argv[2] if len(sys.argv) > 2 else "jobs in Chicago, IL"

response = requests.get(URL, timeout=30)
body = response.text
title = re.search(r"<title>(.*?)</title>", body, re.S)

print("status:", response.status_code)
print("server:", response.headers.get("server"))
print("cf-ray present:", "cf-ray" in response.headers)
print("cf-mitigated:", response.headers.get("cf-mitigated"))
print("title:", title.group(1).strip() if title else None)
print("challenge script in body:", "/cdn-cgi/challenge-platform/" in body)
print("real page marker found:", MARKER in body)

We ran it on September 28, 2026 against Indeed, then against a Stack Overflow question. Both are Cloudflare targets in the benchmark, and each marker is the text the benchmark checks for:

$ python plain_requests.py
status: 403
server: cloudflare
cf-ray present: True
cf-mitigated: None
title: Blocked - Indeed.com
challenge script in body: True
real page marker found: False

$ python plain_requests.py https://stackoverflow.com/questions/11227809/... "Why is conditional processing of a sorted array faster"
status: 403
server: cloudflare
cf-ray present: True
cf-mitigated: challenge
title: Just a moment...
challenge script in body: True
real page marker found: False

Both are HTTP 403 from Cloudflare, but they are different failures. Indeed returned a block page with no cf-mitigated header: the request failed on who it was. Stack Overflow returned cf-mitigated: challenge and the "Just a moment..." page: Cloudflare wanted a browser to run a JavaScript check. The first can be fixed at the connection layer. The second cannot.

What you see What it means What to change
403, server: cloudflare, title "Blocked" or "Attention Required", a Ray ID, no cf-mitigated A WAF rule or a low bot score blocked the request The fingerprint or the IP. Retrying from the same identity does not help
403 (sometimes 503), title "Just a moment...", cf-mitigated: challenge A managed or JavaScript challenge A real browser that can run the challenge, or an API that runs one
200 with a challenges.cloudflare.com script, a data-sitekey, and none of the page's content A Turnstile widget in front of the page Same as above. Detect it by body markup, not by status
Error 1020, "Access denied" A firewall rule the site owner wrote matched your request Usually IP, country or header based; change egress and headers
Error 1015, "You are being rate limited" The site's rate-limiting rule fired Slow down and spread requests across IPs
Error 1010 The owner banned your browser signature Your client fingerprint; move up the ladder
Errors 1006, 1007, 1008 Your IP address is banned Change egress, ideally to residential
Error 1009 Access denied for your country Egress from an allowed country
429 Rate limited by the site or by Cloudflare Lower concurrency per IP, add jitter

Cloudflare documents the 1xxx codes in its troubleshooting reference. The cf-mitigated: challenge header is the most useful single signal: our Stack Overflow request above carried it, and the Indeed block did not.

What people running their own bypass report

The best argument against paying anyone is that the DIY route works. In February 2026 a developer on r/webscraping who could not get past a challenge was told by another, u/HLCYSWAP: "you can't get a 200 with curl_cffi alone because cloudflare only gives the cf_clearance cookie to a real browser that passes the challenge. so you use the browser once to get the cookie, then replay with curl_cffi." curl_cffi's original author, u/yifeikong, replied in the same thread that this is "a pattern" (thread). That is rung 3 below, and it works.

The cost is upkeep. The developer who asked answered a day later: "which works ok until the next challenge comes, so its semi automatic only". In May 2026 another developer wrote that curl_cffi "worked for one year, but today, all of a sudden, it started to get blocked with 403 status codes," across every impersonation profile; the curl_cffi maintainer replied that releasing new fingerprint presets "is kind of a heavier work" (thread). And from the other side of the wall, a developer who runs a Cloudflare-protected site wrote in June 2026: "Clearance lifetime is configurable. We run 30min. That is back to Interactive Challenge every timespan, regardless of prior grant" (thread).

The DIY ladder

What still works in 2026, in order of effort, with the point where each rung stops.

1. Send a real browser's TLS and HTTP/2 fingerprint

This is the cheapest real fix and the one most people skip. curl_cffi binds Python to curl-impersonate, a curl build that reproduces a real browser's TLS ClientHello, HTTP/2 settings and header order. It is a near drop-in for requests.

"""Rung 1: the same request with a real Chrome TLS and HTTP/2 fingerprint.

    pip install curl_cffi
    python curl_cffi_fix.py
"""
import re
import sys

from curl_cffi import requests

URL = sys.argv[1] if len(sys.argv) > 1 else "https://www.indeed.com/l-chicago,-il-jobs.html"
MARKER = sys.argv[2] if len(sys.argv) > 2 else "jobs in Chicago, IL"

# "chrome" resolves to the newest Chrome profile this curl_cffi version ships.
response = requests.get(URL, impersonate="chrome", timeout=30)
body = response.text
title = re.search(r"<title>(.*?)</title>", body, re.S)

print("status:", response.status_code)
print("cf-mitigated:", response.headers.get("cf-mitigated"))
print("title:", title.group(1).strip()[:80] if title else None)
print("bytes:", len(body))
print("real page marker found:", MARKER in body)

With curl_cffi 0.16.3 on September 28, 2026, on the same two pages:

$ python curl_cffi_fix.py
status: 200
cf-mitigated: None
title: 161,000 Jobs, Employment in Chicago, IL | Indeed
bytes: 1059526
real page marker found: True

$ python curl_cffi_fix.py https://stackoverflow.com/questions/11227809/... "Why is conditional processing of a sorted array faster"
status: 403
cf-mitigated: challenge
title: Just a moment...
bytes: 5608
real page marker found: False

The fingerprint alone turned Indeed from a block into the real page. It did nothing for Stack Overflow, because a JavaScript challenge needs a JavaScript engine. We ran both clients against eight of the benchmark's Cloudflare targets, from one laptop connection in Europe, with each target's own URL and marker:

$ python ladder_probe.py
site               requests         curl_cffi
indeed.com         403 blocked      200 page
glassdoor.com      403 blocked      200 page
capterra.com       403 blocked      200 page
stockx.com         403 blocked      403 challenge
crunchbase.com     403 blocked      403 challenge
ziprecruiter.com   403 challenge    403 challenge
stackoverflow.com  403 challenge    403 challenge
axs.com            403 challenge    403 challenge

Three of eight opened. That is the honest size of the TLS fix: real, free, and partial. Two results are worth a second look. On StockX and Crunchbase, fixing the fingerprint moved Cloudflare from a block to a challenge. And Glassdoor and Capterra, which a free library opened from a laptop, returned nothing on all five attempts for three and two of the paid APIs in the benchmark. A paid API is not automatically better than curl_cffi on a given site. Test yours.

Two details matter. First, version drift: impersonate="chrome" resolves to the newest profile your installed curl_cffi ships (chrome150 in 0.16.3), which is what you want; a pinned profile from an old tutorial sends a fingerprint real users stopped sending. Second, consistency: if you set your own User-Agent or sec-ch-ua headers, they must name the same Chrome version as the TLS profile. A mismatch is exactly what the machine-learning engine is trained on.

2. Egress from residential or mobile IPs

Datacenter ranges (AWS, GCP, Hetzner, wherever your scraper runs) start with a worse reputation. Residential and mobile IPs borrow consumer connections and start neutral. Proxies fix only the IP layer: a clean residential IP with a Python handshake is still a bot, and a browser behind a proxy can still leak its real address through WebRTC (camoufox#672 documents a site that checks this). If you drive a browser, turn WebRTC off or route it.

Rotate per session, not per request. Cloudflare scores a visitor across requests. A visitor whose IP changes on every request has no consistent history and looks worse, not better.

3. Hold sessions, and reuse cf_clearance from the same identity

When you pass a challenge, Cloudflare sets a cf_clearance cookie. It lasts for the site's Challenge Passage time, 30 minutes by default and settable from 5 minutes to a year. Cloudflare describes it as tied to the specific visitor and device; in practice, developers report it works only from the same IP, User-Agent and TLS fingerprint that earned it. This is the most common "but I have the cookie" bug. Someone solves the challenge in a real browser, pastes cf_clearance into requests or axios, and is challenged again, because the handshake changed.

The fix is to send later requests from the same identity: the same egress IP, the same User-Agent string, and curl_cffi with the impersonate profile that matches the browser that solved it. Practical shape: one cookie jar per proxy IP, a page or two of warm-up before the deep URLs, and a re-solve in the browser tier only when the clearance expires. Treat cf-mitigated: challenge on a warm session as "re-solve", not "retry".

4. Use a stealth browser for JavaScript challenges

When a site serves a JavaScript challenge, you need a real browser engine that does not look automated. Stock Playwright, Puppeteer and Selenium leak automation markers at the protocol level, and Cloudflare looks for them.

nodriver drives Chrome directly over the DevTools Protocol, with no WebDriver and no Playwright layer. zendriver is its community fork with a faster release cadence. In a third-party 2026 benchmark of seven stealth tools across 31 Cloudflare targets, nodriver passed the most (Ian Paterson, dev.to): 28 of 31 targets, against 26 for curl_cffi, 25 for Patchright and Camoufox, and 24 for stock Playwright. Treat those as one tester's results on one day.

Camoufox is a Firefox fork that applies its fingerprint changes inside the browser's C++ code rather than by injecting JavaScript, so page scripts cannot see the patches. Its maintainers moved it to a Firefox 152 base in July 2026. That is the cost of this rung: a full browser per session is slow and memory-hungry next to an HTTP client, and at volume you run a fleet of them.

Skip FlareSolverr for new work. It runs Selenium with undetected-chromedriver behind a small HTTP service, and on many current Cloudflare configurations it clicks the challenge in a loop until it times out. Results vary by site: some users still report undetected-chromedriver working, and vendors publish opposite claims about it. Its own tracker documents the loop (#1675, #1732), and a maintainer wrote in April 2026 that there are "a tonne of identical issues". If FlareSolverr sits in a pipeline you own, it is a likely source of intermittent timeouts.

Two cautions for this whole rung. Cloudflare can detect DevTools Protocol automation itself, so no CDP-driven browser is invisible. And these are open-source tools in a contest that does not end: Crawlee shipped v3.18.0 to handle a change in Cloudflare's challenge page markup (release notes), and anyone on an older version failed those sites without an error. You inherit that maintenance.

5. Pace and warm up like a person

Rate limits are the easy part: slow down, spread across IPs, and a 429 or a 1015 clears. Session scoring is the hard part. A session that lands on a deep product URL in 40 milliseconds with no scroll, no pointer events and no time on page scores badly however good its fingerprint is. If you run a browser, add real dwell time, a scroll, and a path (home, category, item). If you run an HTTP client, keep per-IP concurrency at one, space requests by seconds with jitter, and accept that some sites will not open to a client that cannot run their JavaScript.

What stopped working

  • A browser User-Agent pasted onto requests. Cloudflare compares it with the TLS handshake and HTTP/2 shape; a User-Agent alone is a mismatch.
  • cloudscraper and other Python solvers for the JavaScript challenge. The challenge script changes faster than the libraries.
  • puppeteer-extra-plugin-stealth. Its patches have a recognizable fingerprint of their own.
  • FlareSolverr, for the reasons above.
  • Copying cf_clearance between clients or IPs.
  • Rotating IPs on every request on sites that score sessions.
  • Any tool that has not shipped a release since Cloudflare's last challenge change.

One method we do not cover: finding the site's origin server IP and sending requests to it directly, around Cloudflare. It works only when the site owner leaked infrastructure they did not intend to expose, and it takes you from reading a public page to reaching a server the owner chose to hide. We do not recommend it.

Getting a page is not the same as completing an action

Everything above is about retrieving a public page: a job listing, a product page, a question and its answers. That is what the benchmark measures, and it is what most scraping is. A different job is completing a protected action: logging in, submitting a form, adding to a cart, checking out. Sites put Turnstile on those actions on purpose.

The mechanics make the difference plain. A Turnstile widget produces a token that the site's server must check with Cloudflare, and each token can be validated only once and expires 300 seconds after it is issued (Cloudflare docs: server-side validation). A token is tied to the widget's sitekey and the hostnames it was set up for. You cannot collect tokens in advance, spend one twice, or use one from a different session. CAPTCHA-solving services work by running the widget for you and handing back a token, which you then have to submit from a matching session inside that window.

Two practical points follow. First, if your target is a public page and you keep hitting Turnstile, the widget is usually an interstitial in front of the page, and the fixes in the ladder apply: a clean identity in a real browser passes Non-Interactive and Invisible widgets with no click. Detect it by markup, not by status: a 200 whose body carries the challenges.cloudflare.com script and a data-sitekey, and none of the page's content, is a gate. We once had a job report zero successes for two months because a target served Turnstile from its own domain at HTTP 200 and our detector only looked at Cloudflare's headers. Second, if your job is the action itself, on an account or a form that is not yours, stop and get permission or an API from the site. String's fetch API is built for public pages, and we do not sell it for the second job.

How to test a Cloudflare bypass before you trust it

Whatever you build or buy, measure it on your own URLs before you trust it. This is the method our benchmark uses, cut down to what one team can run in an afternoon.

  1. Pick 50 or more real URLs from the sites you need, in the proportions you will fetch them. Include your hardest site. One URL per site hides everything that matters.
  2. Write a content check per URL. A phrase, a price selector or a field that only the real page carries. Status codes lie: Turnstile at 200 and an empty shell at 200 both pass a status check.
  3. Run each URL at least five times, spread over a day, with any response caching off on your side and on the vendor's.
  4. Count challenge leakage. Log every response that contains Just a moment, /cdn-cgi/challenge-platform/, challenges.cloudflare.com or cf-mitigated: challenge. A vendor that returns challenge pages as successes will look good on a status dashboard and poison your data.
  5. Record latency on successes and on failures. A 90-second timeout is a cost too.
  6. Compute cost per valid page, not per request: total spend divided by requests that passed step 2. Include retries you had to make.
  7. Repeat monthly. Cloudflare and the vendors both change. Our own Cloudflare column moved by more than 80 points for one vendor between August and September.

If you want the full version, the benchmark harness does all of this for 16 APIs and takes a custom target list.

The managed route: String's fetch API

At some point the engineering time you spend on the ladder costs more than an API that turns "get me this page" into one call. This is what we sell, so read this section as a vendor claim with public tests attached.

String's Web Access API takes a URL and chooses the fetch path itself: a plain request, a real browser, residential proxies and challenge handling, only when the page needs them (docs). Here is the Python call on the same Indeed page and the Stack Overflow page that stopped curl_cffi:

"""Fetch a Cloudflare-protected page through the String Web Access API.

    pip install requests
    export STRING_API_KEY=...
    python string_fetch.py https://www.indeed.com/l-chicago,-il-jobs.html "jobs in Chicago, IL"
"""
import os
import sys
import time

import requests

url = sys.argv[1] if len(sys.argv) > 1 else "https://www.indeed.com/l-chicago,-il-jobs.html"
marker = sys.argv[2] if len(sys.argv) > 2 else "jobs in Chicago, IL"

started = time.time()
response = requests.post(
    "https://request.usestring.ai/v1/fetch",
    headers={"Authorization": f"Bearer {os.environ['STRING_API_KEY']}"},
    json={"url": url, "format": "markdown"},
    timeout=120,
)
elapsed = time.time() - started

print("api status:", response.status_code)
print("site status:", response.headers.get("x-status-code"))
print("billed as:", response.headers.get("x-billed-request-type"))
print("markdown characters:", len(response.text))
print("real page marker found:", marker in response.text)
print(f"seconds: {elapsed:.1f}")

What it printed on September 28, 2026:

$ python string_fetch.py https://www.indeed.com/l-chicago,-il-jobs.html "jobs in Chicago, IL"
api status: 200
site status: 200
billed as: request_standard
markdown characters: 55973
real page marker found: True
seconds: 2.3

$ python string_fetch.py https://stackoverflow.com/questions/11227809/... "Why is conditional processing of a sorted array faster"
api status: 200
site status: 200
billed as: request_premium
markdown characters: 394216
real page marker found: True
seconds: 7.9

The x-billed-request-type header tells you which path the page needed, and so what it cost. Indeed opened on the cheapest path, a plain request on a standard proxy. Stack Overflow needed the premium proxy class. x-status-code is the site's own status, so you can still tell a real 404 from a block.

The same call in TypeScript, with no dependencies (Node 18 or later, or Bun):

// Fetch a Cloudflare-protected page through the String Web Access API.
//   export STRING_API_KEY=...
//   npx tsx string_fetch.ts <url> "<text the real page contains>"   (or: bun string_fetch.ts ...)

const [
  url = "https://stackoverflow.com/questions/11227809/why-is-processing-a-sorted-array-faster-than-processing-an-unsorted-array",
  marker = "Why is conditional processing of a sorted array faster",
] = process.argv.slice(2);

const started = Date.now();
const response = await fetch("https://request.usestring.ai/v1/fetch", {
  method: "POST",
  headers: {
    Authorization: `Bearer ${process.env.STRING_API_KEY}`,
    "Content-Type": "application/json",
  },
  body: JSON.stringify({ url, format: "markdown" }),
});
const markdown = await response.text();

console.log("api status:", response.status);
console.log("site status:", response.headers.get("x-status-code"));
console.log("billed as:", response.headers.get("x-billed-request-type"));
console.log("markdown characters:", markdown.length);
console.log("real page marker found:", markdown.includes(marker));
console.log(`seconds: ${((Date.now() - started) / 1000).toFixed(1)}`);

Run with Bun 1.3.14 on September 28, 2026:

$ bun string_fetch.ts
api status: 200
site status: 200
billed as: request_premium
markdown characters: 422041
real page marker found: true
seconds: 14.2

$ bun string_fetch.ts https://www.indeed.com/l-chicago,-il-jobs.html "jobs in Chicago, IL"
api status: 200
site status: 200
billed as: request_standard
markdown characters: 56001
real page marker found: true
seconds: 2.2

The Stack Overflow call took 14.2 seconds here against 7.9 seconds from Python an hour earlier; the page and path were the same, so treat single timings as a range, not a promise. Then the same eight sites from the ladder, through String:

$ python string_ladder.py
site               result     billed as          seconds
indeed.com         page       request_standard   2.2
glassdoor.com      page       request_premium    2.0
capterra.com       page       request_premium    3.8
stockx.com         page       request_premium    8.4
crunchbase.com     page       request_premium    5.6
ziprecruiter.com   page       request_standard   2.1
stackoverflow.com  page       request_premium    4.9
axs.com            page       browser_premium    61.2

Eight of eight, with one attempt each. AXS, the hardest Cloudflare site in the benchmark, needed a real browser on a residential proxy and took 61 seconds. markdown is one of three output formats; json (the default) wraps the site's status, headers and body, and raw returns the original bytes. If you work in an agent, the same fetch is a tool call through String's MCP server.

Limits, stated plainly. String fetches the pages you point it at; whole-site crawling is your loop to write, with /sitemap for URL discovery. It is not self-hostable. It is for public pages, not for logging in or submitting forms on a site that has not agreed to it. Some destinations need approval or KYC before String will fetch them (access control). And it does not open everything: in the September 16 run, String returned 485 of 500 requests, and temu.com returned nothing to any of the 16 APIs.

What it costs

String bills per successful request, at a rate set by the path the page needed. From the pricing page, per 1,000 requests:

Path Starter ($20 a month) Growth ($100 a month)
Plain request, standard proxy $0.30 $0.20
Plain request, premium proxy $3.00 $2.00
Browser, standard proxy $1.50 $1.00
Browser, premium proxy $6.00 $4.00

The first 5,000 standard requests are free. A blocked request is not billed. There are no credit multipliers and no per-GB charge on fetch.

Worked through on the eight-site run above: two pages billed at the standard request rate, five at premium request, one at premium browser. At that mix, 1,000 pages cost about $21.60 on Starter and $14.40 on Growth, before the plan fee. Your mix depends on your sites, and it can change when a site changes its Cloudflare settings. The header on every response tells you which rate applied, so you can measure it on your own URLs before you commit.

For comparison, the DIY route's costs are residential bandwidth (sold per GB by proxy vendors), servers for a browser fleet if you need rung 4, and the engineering time to keep up with Cloudflare. We do not compare other APIs' prices on this page; their rate cards bill by credits, by GB or by difficulty tier, and a like-for-like number needs your own URL list.

Which route to use

At hobby scale, a few thousand pages from sites that only check the connection, do it yourself. Start with curl_cffi and a small pool of residential IPs, and add a stealth browser only for the sites that serve a JavaScript challenge. You will learn how the layers work.

At production scale, where a broken scraper means missing data that someone downstream counts on, buy the managed route. The deciding cost is engineering time, not the per-request price. The day your success rate drops after a Cloudflare change and someone stops their real work to fix it, you have found the line.

The middle case, steady volume on a stable list of medium-hard sites, is a judgment call. If the list is stable, curl_cffi plus a stealth browser plus good proxies will hold. If the list keeps growing or the sites keep hardening, the API wins. Either way, run the test in the section above on your own URLs first. For blocks that are not Cloudflare (rate limits, proxies, sessions, soft blocks at HTTP 200), see how to scrape a website without getting blocked. For the wider field beyond Cloudflare, see the best web scraping APIs in 2026.

FAQ

How do I bypass Cloudflare when scraping?

Read the response first. A 403 "Blocked" page means the fingerprint or IP failed: use curl_cffi with impersonate="chrome" and a residential IP. A 403 "Just a moment..." page with cf-mitigated: challenge needs a real browser such as nodriver or Camoufox, or a fetch API that runs one. On eight benchmark Cloudflare sites on September 28, 2026, curl_cffi opened three and String's API opened eight.

How to get past Cloudflare Turnstile when scraping?

On a public page, Turnstile is usually Non-Interactive or Invisible, and a real browser with a clean identity on a residential IP passes it without a click. An HTTP client cannot. Detect it by body markup (challenges.cloudflare.com, data-sitekey), because it can arrive at HTTP 200. On a login or checkout form, the token is single-use and expires after 300 seconds; that is an action, not a page, and it needs the site's permission.

Which scraping service works on Cloudflare-protected websites?

In the September 16, 2026 run of the open Web Data Frontier Benchmark, on 15 Cloudflare sites, String passed 100% of requests, Context.dev 96.0%, Firecrawl and Apify 93.3%, ScrapingBee 92.0%, Scrapfly 90.7%, ScraperAPI 89.3% and Bright Data 88.0%. Five of the 16 APIs passed under 50%. Test your own URLs; the ranking moves between runs.

How does Cloudflare Turnstile detect bots, and can scraping APIs beat it?

Turnstile runs checks in the browser (environment, automation markers, and proof-of-work style challenges) and combines them with Cloudflare's IP and fingerprint signals, then issues a token the site verifies server-side. APIs that run a real browser with clean residential egress pass it on public pages. No API can reuse or pre-mint tokens, because each is bound to one sitekey, one session and 300 seconds.

How to bypass TLS / JA3 fingerprinting when scraping?

Send a real browser's handshake. In Python, curl_cffi with impersonate="chrome" reproduces Chrome's TLS ClientHello and HTTP/2 settings; in our test it turned Indeed from a 403 into the real page. Keep your User-Agent on the same Chrome version as the profile. Cloudflare also computes JA4, which sorts extensions, so random extension shuffling does not help.

Can I reuse a cf_clearance cookie?

Yes, but only from the identity that earned it: the same IP, User-Agent and TLS fingerprint, for the challenge passage time the site set. Pasting it into requests from a browser session fails, because the handshake changed. Solve once in a browser, then send follow-up requests with curl_cffi using the matching profile through the same proxy IP.

Are open-source anti-bot solvers reliable enough to run in production, or do I need a paid API?

Some are, for some sites. curl_cffi is stable and opened three of eight Cloudflare sites for us; nodriver and Camoufox handle JavaScript challenges at the cost of a browser fleet. FlareSolverr is not reliable on current Cloudflare. The deciding cost is maintenance: every Cloudflare change lands on your team. Paid APIs varied from 8% to 100% on the benchmark's Cloudflare sites, so a paid API is not automatically reliable either.

Is it legal to bypass Cloudflare to scrape a website?

Getting past a bot check to read a public page is a technical act, not automatically an illegal one, but it depends on what you collect, the site's terms and where you operate. Logging in, submitting forms, or collecting personal data changes the analysis. This is not legal advice; see is web scraping legal and does robots.txt legally prevent scraping.

Sources

Cheers,
String team

Get your API key →Explore the Web Access API
© 2026 StringEU and UK GDPR Article 27 representative — appointment verified by EuverifyBuilt in New York City 🗽 🍎