A scraper gets a 403 Forbidden when the site understood the request and refused it, almost always because a bot check scored the request as automated. You fix it by finding which signal failed and changing that signal, cheapest fix first.
On October 7, 2026 we sent a default Python requests call to the 100 bot-protected pages in String's benchmark, and 51 answered 403. A full set of Chrome headers turned 17 of the 51 into the real page. Chrome's TLS fingerprint fixed 8 more. 26 stayed blocked, and those need a clean IP, a real browser or a scraping API.
The 403 fix ladder ran three rungs on each of the 100 pages in the Web Data Frontier Benchmark, one request per page per rung, from a US home connection with no proxy.
Rung 1 was requests with its default headers. Rung 2 was requests with the 13 headers Chrome 150 sends when it opens a page. Rung 3 was curl_cffi with impersonate="chrome", which sends the same headers over Chrome's TLS and HTTP/2 fingerprint. So rung 2 changed only the headers, and rung 3 changed only the fingerprint.
A page counted as fixed only when it came back with a 2xx status and the page-specific text the benchmark checks for. A 200 without that text did not count.
| Anti-bot vendor (benchmark label) | 403 on rung 1 | Fixed by headers | Fixed by TLS fingerprint | Still blocked |
|---|---|---|---|---|
| DataDome | 16 | 4 | 3 | 9 |
| Cloudflare | 15 | 7 | 2 | 6 |
| Akamai | 9 | 1 | 1 | 7 |
| PerimeterX | 5 | 2 | 0 | 3 |
| AWS WAF | 2 | 0 | 1 | 1 |
| In-house or other | 4 | 3 | 1 | 0 |
| All | 51 | 17 | 8 | 26 |
Four results stand out.
Across all 100 pages, real content rose from 28 on rung 1 to 48 on rung 2 and 61 on rung 3. Treat this as a snapshot: one request per page, one day, one home IP. Capterra opened on rung 2 and refused rung 3 about a minute later.
Try one URL that blocks you now: curl -X POST https://request.usestring.ai/v1/fetch -H "Authorization: Bearer $STRING_API_KEY" -d '{"url":"<page>"}'. The first 5,000 standard requests are free.
RFC 9110 defines 403 as a server that "understood the request but refuses to fulfill it", and it lets the server explain why in the response body (RFC 9110, section 15.5.4). So print response.text and the response headers first. The body names the wall more often than not.
| What the 403 body or headers show | Likely cause | First fix |
|---|---|---|
"Just a moment..." or a cf-mitigated: challenge header |
Cloudflare JavaScript challenge | A real browser, or a scraping API |
| "Attention Required! | Cloudflare" or "Sorry, you have been blocked" | Cloudflare block rule: headers, fingerprint or IP | Full Chrome headers, then curl_cffi |
A small page that loads captcha-delivery.com |
DataDome device check | curl_cffi, then a clean residential IP |
| "Access Denied" with a "Reference #" | Akamai edge block | Chrome fingerprint, residential IP, browser for the sensor |
| "Access to this page has been denied" | PerimeterX (HUMAN) | Full headers, then a browser |
| A 403 on a JSON or API URL that works in your browser tab | Missing cookie, CSRF token or Referer | Copy the browser's cookies and Referer into a session |
| A 403 after many successful requests | IP reputation or pace | Slow down, rotate the IP per session |
| A 403 with a login or "permission" message | A real permission check | Log in with an account you own, or stop |
Cloudflare documents the cf-mitigated: challenge header as the signal on every challenge page (Cloudflare docs). The last row is the only 403 that is not a bot check. Respect it.
String is the most reliable web scraping API, bypassing anti-bot, and you only pay when content comes back. Here is what that means for a 403, one fact per line:
POST /v1/fetch with a URL, and String picks the path the page needs: a plain request, a premium residential proxy or a real browser (fetch docs).x-billed-request-type header names it, such as request_standard, request_premium or browser_premium, and x-status-code carries the site's own status, so a real 404 stays a 404.Each step below costs more than the one before it, so stop at the first one that returns the real page.
A default requests call sends four headers and says python-requests/2.34.2 in its user agent. Chrome sends 13 or more when it opens a page, including Sec-CH-UA client hints and the Sec-Fetch-* family. Copy the whole set from one browser version, not just the user agent. A Chrome user agent with no client hints is its own red flag.
This is the fix in the most-viewed Stack Overflow question on the topic, asked in 2013 and viewed 356,747 times. The top answer, with 378 votes, says to try "setting a known browser user agent" (Stack Overflow). In our run that class of fix still opened a third of the 403s (17 of 51).
When headers do not help, the site is reading your handshake. Python requests negotiates TLS like OpenSSL, not like Chrome, and it speaks only HTTP/1.1, so the site sees a Chrome user agent on a non-Chrome connection. One r/webscraping user described this exactly:
"why, if the headers are identical to my browser, and its coming from a trustworthy ip, do all my requests get hit with a 403?"
u/Mugwartz, r/webscraping, May 2024
The same user reported a 200 after switching to curl_cffi with the impersonate argument. In our run that switch opened 8 more of the 51 pages.
If the fingerprint matches and you still get 403, the site is scoring your IP or your session. Datacenter and cloud addresses score worse than home and mobile ones, which is why a scraper can work on a laptop and fail on a server; our page on scrapers blocked in Docker or Lambda covers that case.
Keep cookies in a requests.Session, carry a Referer for API calls, and rotate the IP per session rather than per request. A 403 that starts after many good requests usually means pace; handling rate limits covers the backoff.
A "Just a moment..." page, a DataDome device check or an Akamai sensor needs JavaScript to run, and no HTTP client runs it. That is where all 26 of our still-blocked pages sat.
Your options are a stealth browser you maintain, or a scraping API that picks the browser, IP and fingerprint per request. Our guides to bypassing Cloudflare and getting past Akamai Bot Manager go deeper on each vendor. To test the API path on your own failing URLs, run the script below with a free key.
This script tries the free rungs first and sends a URL to String's fetch endpoint only when both fail. It judges every response by the text the real page must contain, so a 200 block page does not pass.
import os
import re
import sys
from typing import Optional
import requests
from curl_cffi import requests as cffi_requests
API = "https://request.usestring.ai/v1/fetch"
API_HEADERS = {"Authorization": f"Bearer {os.environ['STRING_API_KEY']}"}
# The headers Chrome 150 sends when you open a page. Copy them whole:
# a Chrome user agent with no Sec-CH-UA or Sec-Fetch headers looks wrong.
CHROME_HEADERS = {
"Sec-Ch-Ua": '"Not;A=Brand";v="8", "Chromium";v="150", "Google Chrome";v="150"',
"Sec-Ch-Ua-Mobile": "?0",
"Sec-Ch-Ua-Platform": '"macOS"',
"Upgrade-Insecure-Requests": "1",
"User-Agent": ("Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 "
"(KHTML, like Gecko) Chrome/150.0.0.0 Safari/537.36"),
"Accept": ("text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,"
"image/webp,image/apng,*/*;q=0.8"),
"Sec-Fetch-Site": "none",
"Sec-Fetch-Mode": "navigate",
"Sec-Fetch-User": "?1",
"Sec-Fetch-Dest": "document",
"Accept-Encoding": "gzip, deflate",
"Accept-Language": "en-US,en;q=0.9",
}
# Text that a vendor's block or challenge page carries, and a real page does not.
WALLS = {
"Cloudflare challenge": r"Just a moment\.\.\.|cf_chl_|challenges\.cloudflare\.com",
"Cloudflare block": r"Attention Required! \| Cloudflare|Sorry, you have been blocked",
"DataDome": r"captcha-delivery\.com",
"PerimeterX": r"px-captcha|Access to this page has been denied",
"Akamai": r"Access Denied.{0,400}Reference #|errors\.edgesuite\.net|bm-verify|sec-if-cpt",
"AWS WAF": r"awswaf|gokuProps",
"Kasada": r"KPSDK",
}
def diagnose(status: int, headers: dict, html: str, must_contain: str) -> Optional[str]:
"""Return None for a real page, or the name of what blocked you."""
if must_contain.lower() in html.lower() and 200 <= status < 300:
return None
if headers.get("cf-mitigated") == "challenge":
return "Cloudflare challenge"
for wall, pattern in WALLS.items():
if re.search(pattern, html[:300_000], re.I | re.S):
return wall
return f"HTTP {status}, page text missing"
def fetch(url: str, must_contain: str) -> dict:
# Rung 1: plain requests with a full Chrome header set. Free, and fixes header checks.
r = requests.get(url, headers=CHROME_HEADERS, timeout=25)
wall = diagnose(r.status_code, r.headers, r.text, must_contain)
if wall is None:
return {"url": url, "fixed_by": "headers", "status": r.status_code}
# Rung 2: same headers, Chrome's TLS and HTTP/2 fingerprint. Fixes fingerprint checks.
c = cffi_requests.get(url, impersonate="chrome", timeout=25)
wall2 = diagnose(c.status_code, c.headers, c.text, must_contain)
if wall2 is None:
return {"url": url, "fixed_by": "tls", "status": c.status_code, "first_wall": wall}
# Rung 3: the page needs a clean IP, a browser or a solved challenge. Send it to the API.
api = requests.post(API, headers=API_HEADERS, json={"url": url}, timeout=120)
api.raise_for_status()
body = api.json()
html, status = str(body.get("data", "")), body.get("statusCode") or 0
still = diagnose(status, {}, html, must_contain) # check the API result the same way
return {"url": url, "fixed_by": "api" if still is None else None, "status": status,
"first_wall": wall, "second_wall": wall2, "still_blocked": still,
"bytes": len(html)}
if __name__ == "__main__":
# usage: python fix_403.py URL "text the real page contains" [URL "text" ...]
args = sys.argv[1:]
for url, text in zip(args[::2], args[1::2]):
print(fetch(url, text))
We ran it on October 7, 2026 with pip install requests curl_cffi. A static test page passed on rung 1. Three pages that refused every free rung that morning came back through the API as full pages with the expected text: a Stack Overflow question behind a Cloudflare challenge (1.1 MB), a Home Depot product behind Akamai (789 KB) and a G2 category page behind DataDome (538 KB).
Pick must_contain with care: a product name, a listing ID or a price field from the page you expect. A short, common word can also appear on a block page.
The most expensive 403 is the one your code thinks it fixed. In our run, six Akamai pages answered the header or fingerprint fix with HTTP 200 and about 2.5 KB of script, and no product. LinkedIn answered the default client with status 999, and Walmart answered it with a 200 "Robot or human?" page. Status-code checks pass all of those.
Decide success by content: the text you expect, the item count, the byte range of a known-good page. Our answer on detecting a bot block behind HTTP 200 covers the checks, and getting past CAPTCHA covers the pages that end in a puzzle.
Many 403s need nothing more than Step 1. One r/webscraping poster who could not get past a 403 was told to read the body first:
"get the response.text to see what it says"
u/RHiNDR, r/webscraping, July 2025
The poster found a Cloudflare challenge, copied the headers from the browser's network tab, and reported a 200 the same day. If one site blocks you and a header change fixes it, you do not need an API. The ladder matters when you scrape many protected sites at once, because each one fails at a different rung.
Start with the URL list your scraper fails on today. Run the script above on it with a free key, and read the fixed_by column: it tells you which rung each site needs.
The site's bot check scored your request as automated and refused it. Read the 403 body to see which vendor sent it, then fix the cheapest failing signal first: full browser headers, then a Chrome TLS fingerprint with curl_cffi, then a clean IP or a real browser. On October 7, 2026, headers fixed 17 of 51 blocked benchmark pages and the fingerprint fixed 8 more.
Replace the default requests headers with a full Chrome header set, including the Sec-CH-UA and Sec-Fetch-* headers. If that fails, switch to curl_cffi with impersonate="chrome" so the TLS and HTTP/2 handshake matches the user agent. If the 403 page is a JavaScript challenge, you need a real browser or a scraping API.
When you scrape a public page, almost always yes. RFC 9110 says a 403 means the server understood the request and refused it, and it may say why in the body. A 403 that mentions login or permissions is a real access rule, and a different account or method is the only fix.
Your browser and Postman send a different TLS fingerprint and more headers than Python requests, which also lacks HTTP/2. Bot checks from Cloudflare, Akamai and others compare the fingerprint with the user agent. One r/webscraping user with a valid account fixed this by switching to curl_cffi.
requests raises that error from raise_for_status() when the server answers 403. The message names only the status. Catch the error, print response.text and the response headers, and look for the vendor's block page to learn which fix applies.
PerimeterX (now HUMAN) returns a 403 page titled "Access to this page has been denied" when its sensor scores the request as a bot. In our October 7 run, full Chrome headers opened 2 of the 5 PerimeterX pages that sent a 403, and the TLS fingerprint opened none of the other 3. Those need a real browser.
Only when the IP is the cause. A VPN changes your address, but not your headers or TLS fingerprint, and a VPN exit is often a datacenter address that scores no better. Fix headers and fingerprint first, then move to residential IPs if the 403 persists.
Make every part of the request agree with one real browser: headers, TLS fingerprint, IP type, cookies and pace. Check each response by its content, not its status. Our guide to scraping without getting blocked covers each signal in detail.
requests 2.34.2 and curl_cffi 0.16.3 (Chrome 150 profile). Pass rule and target text from src/check.ts and src/tests.const.ts in the public harness. Per-target results from the September 16, 2026 run.