Handle rate limiting in four moves. First, find out who limited you: the website you are scraping, or the scraping API or proxy you send the request through. Second, cap your request rate per site and, separately, per API key. Third, when you get an HTTP 429, wait for the Retry-After time if the response gives one, and otherwise back off exponentially with random jitter. Fourth, if one site keeps limiting you at a polite rate, change the route, not the speed: a different IP pool or a scraping API that rotates it. In String's Web Data Frontier Benchmark run of September 16, 2026, 187 of 8,000 requests failed on a rate limit, and 159 of those (85%) were the API refusing the client, not the website refusing the API.
A rate limit is a cap on how many requests one client may send in a time window. Past the cap, the server answers 429 Too Many Requests, a challenge page, or a slow or empty response. HTTP defines 429 in RFC 6585, and the Retry-After header in RFC 9110. When you scrape through a service, two different servers can limit you, and the fix is different for each:
| Who limits you | What it looks like | What it counts | The fix |
|---|---|---|---|
| The target website | 429 or a block page from the site, often after a burst | Requests per IP address, session or account | Slow down per site, spread load across IPs, retry later |
| Your scraping API or proxy | 429 from the API itself, or an error message about your plan | Requests per API key, per second or in parallel | Stay under your plan's rate, smooth bursts, ask for a higher limit |
The mistake is to treat both the same way. Slowing your whole scraper down because one site sent a 429 wastes the capacity you pay for on every other site. Rotating IPs because your API key hit its plan limit does nothing, because the key is what is being counted.
The benchmark sends 16 scraping APIs the same 100 bot-protected pages, five attempts each, 8,000 requests in all. The harness is gentle: two requests in parallel per provider, with half a second between requests. We went back to the September 16, 2026 results and sorted every failure that carried a rate-limit signal by who sent it.
The counts come from the public results file in the benchmark repository, official_results/benchmark-2026-09-16T01-03-47-074Z.json, and the harness settings are in src/tests.const.ts.
statusCode and in the X-Status-Code header. If your code only checks the outer status, a site's 429 looks like a success.robots.txt for a Crawl-delay line and honor it. Many threads on one domain at once is a common cause of a first 429.X-RateLimit-Limit-Rps header.Retry-After. It comes either as seconds or as an HTTP date, so parse both. Do not retry before it expires.Retry-After. Wait a random time between zero and 1, 2, 4, 8 seconds on each attempt, capped at about 30 seconds. Without jitter, your threads retry in lockstep and hit the limit again together. AWS's write-up on backoff with jitter shows how the random wait spreads them out.This script fetches a list of URLs through String's fetch endpoint. It spaces calls to the API and to each site separately, honors Retry-After in both forms, backs off with full jitter, and pauses only the site that limited it. We ran it on October 3, 2026 against five URLs, one of them a test URL that always returns 429: the four real pages came back on the first try and the test URL gave up after five attempts, as designed.
"""Fetch a list of URLs through a scraping API without tripping rate limits.
pip install requests
export STRING_API_KEY=...
python polite_fetch.py urls.txt
"""
import email.utils
import os
import random
import sys
import threading
import time
from collections import defaultdict
from concurrent.futures import ThreadPoolExecutor
from typing import Optional
from urllib.parse import urlparse
import requests
API = "https://request.usestring.ai/v1/fetch"
HEADERS = {"Authorization": f"Bearer {os.environ['STRING_API_KEY']}"}
API_RATE = 50 # requests per second, kept under the account limit of 60
PER_DOMAIN_GAP = 1.0 # seconds between two requests to the same site
MAX_TRIES = 5
lock = threading.Lock()
next_api_slot = 0.0
next_domain_slot = defaultdict(float)
def wait_turn(domain: str) -> None:
"""Space calls to the API and, separately, to each target site."""
global next_api_slot
with lock:
now = time.monotonic()
start = max(now, next_api_slot, next_domain_slot[domain])
next_api_slot = start + 1 / API_RATE
next_domain_slot[domain] = start + PER_DOMAIN_GAP
time.sleep(max(0.0, start - now))
def retry_after(headers: dict) -> Optional[float]:
"""Retry-After is either seconds or an HTTP date."""
value = next((v for k, v in headers.items() if k.lower() == "retry-after"), None)
if value is None:
return None
if value.strip().isdigit():
return float(value)
when = email.utils.parsedate_to_datetime(value)
return max(0.0, when.timestamp() - time.time())
def backoff(attempt: int) -> float:
"""Exponential backoff with full jitter: 0..1s, 0..2s, 0..4s, capped at 30s."""
return random.uniform(0, min(30, 2 ** attempt))
def fetch(url: str) -> dict:
domain = urlparse(url).netloc
for attempt in range(MAX_TRIES):
wait_turn(domain)
r = requests.post(API, headers=HEADERS, json={"url": url}, timeout=120)
if r.status_code == 429: # the API's own limit: slow the whole client down
time.sleep(retry_after(r.headers) or backoff(attempt))
continue
r.raise_for_status()
body = r.json()
status = body.get("statusCode")
if status == 429 or status in (502, 503): # the target site's limit, inside a 200 envelope
pause = retry_after(body.get("headers", {})) or backoff(attempt)
with lock:
next_domain_slot[domain] = max(next_domain_slot[domain], time.monotonic() + pause)
continue
return {"url": url, "status": status, "tries": attempt + 1, "bytes": len(str(body.get("data", "")))}
return {"url": url, "status": "gave up", "tries": MAX_TRIES}
if __name__ == "__main__":
urls = [line.strip() for line in open(sys.argv[1]) if line.strip()]
with ThreadPoolExecutor(max_workers=20) as pool:
for result in pool.map(fetch, urls):
print(result)
Same rules, written for Node 18+ or Bun. We ran it on the same five URLs on October 3, 2026 with the same result.
// Fetch a list of URLs through a scraping API without tripping rate limits.
// export STRING_API_KEY=...
// npx tsx polite_fetch.ts urls.txt (or: bun polite_fetch.ts urls.txt)
import { readFileSync } from "node:fs";
const API = "https://request.usestring.ai/v1/fetch";
const API_RATE = 50; // requests per second, kept under the account limit of 60
const PER_DOMAIN_GAP_MS = 1000; // between two requests to the same site
const MAX_TRIES = 5;
let nextApiSlot = 0;
const nextDomainSlot = new Map<string, number>();
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));
async function waitTurn(domain: string) {
const now = Date.now();
const start = Math.max(now, nextApiSlot, nextDomainSlot.get(domain) ?? 0);
nextApiSlot = start + 1000 / API_RATE;
nextDomainSlot.set(domain, start + PER_DOMAIN_GAP_MS);
await sleep(start - now);
}
// Retry-After is either seconds or an HTTP date.
function retryAfterMs(headers: Record<string, string>): number | null {
const key = Object.keys(headers).find((k) => k.toLowerCase() === "retry-after");
if (!key) return null;
const value = headers[key].trim();
if (/^\d+$/.test(value)) return Number(value) * 1000;
return Math.max(0, Date.parse(value) - Date.now());
}
// Exponential backoff with full jitter: 0..1s, 0..2s, 0..4s, capped at 30s.
const backoffMs = (attempt: number) => Math.random() * Math.min(30_000, 1000 * 2 ** attempt);
async function fetchPage(url: string) {
const domain = new URL(url).host;
for (let attempt = 0; attempt < MAX_TRIES; attempt++) {
await waitTurn(domain);
const res = await fetch(API, {
method: "POST",
headers: { Authorization: `Bearer ${process.env.STRING_API_KEY}`, "Content-Type": "application/json" },
body: JSON.stringify({ url }),
});
if (res.status === 429) {
// the API's own limit: slow the whole client down
await sleep(retryAfterMs(Object.fromEntries(res.headers)) ?? backoffMs(attempt));
continue;
}
if (!res.ok) throw new Error(`API error ${res.status}`);
const body = await res.json();
if ([429, 502, 503].includes(body.statusCode)) {
// the target site's limit, inside a 200 envelope: pause this site only
const pause = retryAfterMs(body.headers ?? {}) ?? backoffMs(attempt);
nextDomainSlot.set(domain, Math.max(nextDomainSlot.get(domain) ?? 0, Date.now() + pause));
continue;
}
return { url, status: body.statusCode, tries: attempt + 1, bytes: JSON.stringify(body.data ?? "").length };
}
return { url, status: "gave up", tries: MAX_TRIES };
}
const urls = readFileSync(process.argv[2], "utf8").split("\n").map((s) => s.trim()).filter(Boolean);
for (const result of await Promise.all(urls.map(fetchPage))) console.log(result);
If you use Scrapy, its retry middleware does not read Retry-After on its own. An open issue on the Scrapy repository has asked for it since 2019; the reporter noted that the middleware "doesn't seem to even try to honor the retry-after header at any point." Turn on AUTOTHROTTLE_ENABLED, set DOWNLOAD_DELAY and CONCURRENT_REQUESTS_PER_DOMAIN, and extend the retry middleware to sleep for the header's value (Scrapy AutoThrottle docs).
When you call a scraping API, the API's own response is often 200 even when the site behind it said 429. We sent https://httpbin.org/status/429 through String's fetch endpoint on October 3, 2026: the API answered HTTP 200, with "statusCode": 429 in the JSON body, an empty data field, and X-Status-Code: 429 in the headers. Code that checks only response.ok stores an empty page and moves on. Check the inner status, and check that the body holds the content you expected; checking that scraped data is complete covers the content checks.
No universal safe rate exists, and most sites do not publish theirs. A Stack Overflow user asking for one put the problem plainly: "I don't know the time after which I would get blocked (e.g. 1000 requests per day)." One answer said it "depends on the site you are scraping. Sometimes it is documented somewhere but very likely not." Beginners meet it fast: one Python learner on r/learnpython got "a 'too many requests' error" after a handful of test runs against one stats site, and stayed locked out into the next day.
So measure instead of guessing. Start at one request per second per site, watch the share of 429s and empty bodies, and raise the rate only while that share stays near zero. If a site sends Retry-After, it is telling you its window.
String's Web Access API fetches the page and picks the route for each request, so a site's per-IP limit lands on a pool of IP addresses rather than on your server. It allows 60 requests per second per organization with a burst allowance of 3,600 requests and no cap on parallel requests, on every plan, and Enterprise raises the rate (pricing). The first 5,000 standard requests are free. Standard fetch costs $0.30 per 1,000 successful requests on the $20 Starter plan and $0.20 on the $100 Growth plan, and you pay only when String returns content. It works from Python, TypeScript, any HTTP client, the MCP server and Apify.
It does not lift a site's own limit on a logged-in account or an API key you hold for that site; those limits follow the account, whatever the route. If your targets have no bot protection and a low volume, a plain HTTP client with the pacing above may be all you need. To compare providers on your own URLs first, see how to evaluate a web scraping API, and for the full ranking on the 100 sites, best web scraping APIs.
Find out whether the website or your scraping API sent the limit. Then cap your rate per site and per API key separately, honor the Retry-After header, back off exponentially with jitter when there is no header, and pause only the site that limited you. If a site still limits you at a polite rate, change the IP route instead of slowing down further.
As long as the Retry-After header says, in seconds or until the date it gives. Without the header, wait a random time up to 1 second, then up to 2, 4 and 8 seconds on later attempts, capped near 30 seconds. In our September 16, 2026 run, retries a few seconds apart recovered only 5 of 40 rate-limited provider and site pairs.
Check the status code before you parse the body, read Retry-After with email.utils.parsedate_to_datetime when it is a date, sleep, and retry with a capped exponential backoff and random jitter. Space requests per domain with a lock and a next-allowed time for each domain. The script on this page does all of this with requests and a thread pool.
Send fewer requests to each site, spread them over time, keep a consistent browser-like fingerprint, and honor robots.txt crawl delays. For bot-protected sites, IP reputation matters more than speed, so use residential IPs or a scraping API that chooses the route. Then check every response for real content, because many blocks return HTTP 200.
They spread a website's per-IP limit across many addresses, which helps when a site counts requests per IP. They do not help when the limit is on your account, your session cookie or your scraping API key. In our benchmark, 85% of rate-limit failures were the API limiting the client, which no proxy rotation fixes.
There is no single safe rate. Start at about one request per second per site, honor any Crawl-delay in robots.txt, and lower the rate for small sites. Raise it only while the share of 429s and empty pages stays near zero. Large sites with bot protection often limit per IP well before they limit on volume.
Yes. Each plan caps requests per second, requests in parallel, or both, and the API answers 429 past that cap. String allows 60 requests per second per organization with no cap on parallel requests. In our September 16, 2026 run, one provider's API answered 429 on 148 of 500 requests at only two parallel requests.
It depends on the provider. Some bill every request, some refund failures, and some charge only for requests that return content. String's pricing bills only requests that come back with content, so you do not pay for blocks. Check the billing docs of your provider before you plan retries, because retries multiply the bill on providers that charge per attempt.