The best API for crawling an entire website is the one that does two jobs well on your site: it finds every URL, from the sitemap and from links, inside a page cap and a spend cap you set; and it gets each page's content past the site's bot protection, then tells you which pages it missed. Firecrawl's /crawl, Scrapfly's Crawler API, Bright Data's Crawl API and Apify's Website Content Crawler do both in one job. String runs them as two calls: a /sitemap job lists the URLs, then /fetch returns each page as Markdown. Test on your own site first. When we checked the 100 sites in String's benchmark on October 3, 2026, 38 of the 77 sites that publish a sitemap did not serve the sitemap file itself to a plain HTTP client.
People who search for this want one of three things: a list of every URL on a site, the content of every page, or a vendor. The first two are separate jobs with separate failure modes, and it helps to keep them apart even when one API does both.
| Job | What it does | How it fails |
|---|---|---|
| Discovery | Reads robots.txt and the sitemap files, follows links from page to page, stays inside the scope you set |
Misses pages that no sitemap lists and no link reaches; follows links into calendars and filters forever; gets blocked on the sitemap file |
| Content | Fetches each discovered URL and returns HTML, Markdown or JSON | Gets a block page or an empty 200 instead of the content; hits a rate limit halfway through; renders a page before its JavaScript has loaded |
Discovery costs little per URL. Content is a long series of single-page fetches, so a crawl is only as good as the API's success on one protected page, repeated thousands of times.
String's benchmark tests single-page fetches on 100 bot-protected sites: retail, travel, real estate, jobs, tickets, social. It does not test crawl jobs. On October 1 and 3, 2026 we went back to those 100 sites and checked the first thing a crawler reads: robots.txt and the sitemap it points to.
robots.txt, 1,334 sitemap lines between them. The other 23, including Amazon, LinkedIn, Reddit, Yelp, Indeed and Ticketmaster, declare none. On those, discovery has to follow links.Crawl-delay, from 0.2 to 15 seconds. A crawler that honors it on a 10,000-page site at 10 seconds a page needs more than a day. Our guide to handling rate limits when scraping covers pacing.The site list is in the public benchmark harness, in src/tests.const.ts. For each site we fetched the first sitemap that robots.txt declares.
Developers report the same problems in public issue threads. One Firecrawl user reported a crawl where the status endpoint kept returning scraping "but the completed value never changes" (firecrawl #1679). A Crawl4AI user running the docs' own example wrote: "I cannot scrape more than 2 pages - it gets stuck" (crawl4ai #1017). And on r/webscraping, someone copying a 100-page site said its scripts "mess up the normal wget and httrack downloading apps" (thread).
Checked against each vendor's documentation on October 3, 2026. String's benchmark measures single-page fetches, not crawl jobs, so this table compares documented features and does not rank anyone.
| API | Crawl endpoint | Returns page content | Page cap | Depth control | Billing | Results kept |
|---|---|---|---|---|---|---|
| String | POST /v1/sitemap, then POST /v1/fetch per URL |
URLs from /sitemap; content from /fetch |
maxPages 1 to 10,000 |
maxDepth 1 to 100 |
Quote first; billed per page crawled at your plan's rate; fetch billed per successful request | URL lists stay available after the job |
| Firecrawl | POST /v2/crawl |
Yes, Markdown or JSON | limit, default 10,000 |
maxDiscoveryDepth |
1 credit per page; JSON mode 4 more credits per page | 24 hours |
| Scrapfly | Crawler API | Yes, several formats plus WARC and HAR | page_limit (0 = no limit within plan) |
max_depth |
Sum of the Web Scraping API calls the crawl makes; max_api_credit cap |
Not stated |
| Bright Data | Crawl API | Yes, Markdown, text, HTML or JSON | Not documented publicly | Not documented publicly | $1.50 per 1,000 requests pay as you go | Not stated |
| Apify Website Content Crawler | Actor run | Yes | maxCrawlPages |
Max crawling depth | Compute units, about $0.20 per 1,000 pages over raw HTTP and $0.50 to $5 with a browser, by Apify's estimate | Dataset |
| Zyte API | None; /extract takes one URL |
Per URL | You bring the crawler | You bring the crawler | Per request | Not applicable |
Sources: String sitemap reference, Firecrawl crawl docs, Scrapfly Crawler API, Bright Data Crawl API pricing, Apify Website Content Crawler, Zyte API reference. Bright Data's product page prices the same tiers per 1,000 "records" while its pricing page says "requests"; we quote the pricing page.
If you would rather run the crawler yourself, Scrapy, Crawlee and Crawl4AI handle discovery and queueing for free, and you can point their page fetches at any scraping API.
limit of 10,000 pages and returns a 402 before the crawl starts if the balance is short./docs or /products.POST /v1/sitemap with the start URL, maxPages, maxDepth, an optional pathPrefix, and useSitemap: true to seed from /sitemap.xml as well as links. The crawl stays on the start URL's hostname. Nothing is crawled or charged yet.estimatedPages and estimatedCostUsd, a ceiling for the job. For a 10-page crawl of our own site the quote was $0.053; the job was billed $0.003.POST /v1/sitemap/{jobId}/approve starts the crawl and holds the quote against your balance. budgetUsd caps spend; leave it at the default, which is the quote. When we set it far below the quote ($0.01 against $0.053), the job stopped with token_cap_exceeded after 2 of 10 pages, so set it at or above the quote.completed, token_cap_exceeded, canceled or failed.GET /v1/sitemap/{jobId}/urls returns each URL with its statusCode, depth, parentUrl and sourceType (seed, sitemap or HTML link), up to 5,000 per page.POST /v1/fetch with format: "markdown" for each URL that answered 200. Keep under the account rate of 60 requests per second.We ran this on October 3, 2026 against https://usestring.ai/answers with a 10-page cap: the crawl completed, was billed $0.003, and all 9 pages fetched as Markdown returned 200. Without a pathPrefix the crawl left /answers and covered the rest of the site, which is the default behavior: it stays on the hostname, not the path.
"""Crawl a whole site in two steps: discover every URL, then fetch each page as Markdown.
pip install requests
export STRING_API_KEY=...
python crawl_site.py https://example.com/docs 500
"""
import json
import os
import sys
import time
import requests
API = "https://request.usestring.ai/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['STRING_API_KEY']}"}
MAX_QUOTE_USD = 5.00 # refuse to start a crawl quoted above this
def discover(start_url: str, max_pages: int) -> list:
quote = requests.post(f"{API}/sitemap", headers=HEADERS, timeout=60, json={
"url": start_url, "maxPages": max_pages, "maxDepth": 5, "useSitemap": True,
}).json()
print(f"quote: up to {quote['estimatedPages']} pages, at most ${quote['estimatedCostUsd']}")
if float(quote["estimatedCostUsd"]) > MAX_QUOTE_USD:
sys.exit("quote is above MAX_QUOTE_USD; lower maxPages or raise the cap")
job = quote["jobId"]
requests.post(f"{API}/sitemap/{job}/approve", headers=HEADERS, timeout=60).raise_for_status()
while True:
status = requests.get(f"{API}/sitemap/{job}", headers=HEADERS, timeout=60).json()
if status["status"] not in ("awaiting_approval", "running"):
break
time.sleep(5)
print(f"crawl {status['status']}, cost ${status.get('costUsd')}")
urls, offset = [], 0
while True:
page = requests.get(f"{API}/sitemap/{job}/urls", headers=HEADERS, timeout=60,
params={"limit": 1000, "offset": offset}).json()
urls += page["urls"]
offset += len(page["urls"])
if not page["urls"] or offset >= page["total"]:
return urls
def fetch_markdown(url: str) -> dict:
r = requests.post(f"{API}/fetch", headers=HEADERS, timeout=120, json={"url": url, "format": "markdown"})
if r.status_code == 429: # over the account rate: wait and try once more
time.sleep(2)
r = requests.post(f"{API}/fetch", headers=HEADERS, timeout=120, json={"url": url, "format": "markdown"})
# Markdown comes back as the response body; the site's own status is in X-Status-Code.
return {"url": url, "status": int(r.headers.get("X-Status-Code", r.status_code)), "markdown": r.text}
if __name__ == "__main__":
start, limit = sys.argv[1], int(sys.argv[2])
found = discover(start, limit)
pages = [u["url"] for u in found if u["statusCode"] == 200 and not u["isSitemap"]]
failed = [u for u in found if u["statusCode"] != 200 and not u["isSitemap"]]
print(f"discovered {len(found)} URLs: {len(pages)} answered 200, {len(failed)} did not")
with open("pages.jsonl", "w") as out:
for url in pages:
result = fetch_markdown(url)
out.write(json.dumps(result) + "\n")
print(result["status"], len(result["markdown"]), url)
for u in failed: # keep this list: these pages are missing from your copy
print("MISSING", u["statusCode"], u["url"], "from", u["parentUrl"])
The same flow for Node 18+ or Bun. Same run, same result on October 3, 2026.
// Crawl a whole site in two steps: discover every URL, then fetch each page as Markdown.
// export STRING_API_KEY=...
// npx tsx crawl_site.ts https://example.com/docs 500 (or: bun crawl_site.ts ...)
import { writeFileSync } from "node:fs";
const API = "https://request.usestring.ai/v1";
const HEADERS = { Authorization: `Bearer ${process.env.STRING_API_KEY}`, "Content-Type": "application/json" };
const MAX_QUOTE_USD = 5; // refuse to start a crawl quoted above this
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));
type Found = { url: string; statusCode: number; isSitemap: boolean; parentUrl: string | null };
async function discover(startUrl: string, maxPages: number): Promise<Found[]> {
const quote = await fetch(`${API}/sitemap`, {
method: "POST",
headers: HEADERS,
body: JSON.stringify({ url: startUrl, maxPages, maxDepth: 5, useSitemap: true }),
}).then((r) => r.json());
console.log(`quote: up to ${quote.estimatedPages} pages, at most $${quote.estimatedCostUsd}`);
if (Number(quote.estimatedCostUsd) > MAX_QUOTE_USD) throw new Error("quote is above MAX_QUOTE_USD");
await fetch(`${API}/sitemap/${quote.jobId}/approve`, { method: "POST", headers: HEADERS });
let status;
do {
await sleep(5000);
status = await fetch(`${API}/sitemap/${quote.jobId}`, { headers: HEADERS }).then((r) => r.json());
} while (["awaiting_approval", "running"].includes(status.status));
console.log(`crawl ${status.status}, cost $${status.costUsd}`);
const urls: Found[] = [];
for (let offset = 0; ; ) {
const page = await fetch(`${API}/sitemap/${quote.jobId}/urls?limit=1000&offset=${offset}`, { headers: HEADERS }).then((r) => r.json());
urls.push(...page.urls);
offset += page.urls.length;
if (!page.urls.length || offset >= page.total) return urls;
}
}
async function fetchMarkdown(url: string) {
const res = await fetch(`${API}/fetch`, { method: "POST", headers: HEADERS, body: JSON.stringify({ url, format: "markdown" }) });
// Markdown comes back as the response body; the site's own status is in X-Status-Code.
return { url, status: Number(res.headers.get("x-status-code") ?? res.status), markdown: await res.text() };
}
const [start, limit] = [process.argv[2], Number(process.argv[3])];
const found = await discover(start, limit);
const pages = found.filter((u) => u.statusCode === 200 && !u.isSitemap).map((u) => u.url);
const failed = found.filter((u) => u.statusCode !== 200 && !u.isSitemap);
console.log(`discovered ${found.length} URLs: ${pages.length} answered 200, ${failed.length} did not`);
const results = [];
for (const url of pages) {
const r = await fetchMarkdown(url);
results.push(r);
console.log(r.status, r.markdown.length, url);
}
writeFileSync("pages.jsonl", results.map((r) => JSON.stringify(r)).join("\n") + "\n");
for (const u of failed) console.log("MISSING", u.statusCode, u.url, "from", u.parentUrl); // pages missing from your copy
A crawl job that reports completed has finished its queue. It has not proved it reached every page. Three checks catch most gaps:
/sitemap results carry a status for every URL; keep the list and retry it.String's Web Access API suits teams that want the URL list and the content as separate, inspectable steps: discovery with a quote before anything is charged, then content from the fetch API that passed 485 of 500 requests (97.0%) on the 100 bot-protected sites in our September 16, 2026 benchmark, the highest of 16 providers. We build String and run that benchmark, so test it on your own site.
To pull an archived copy of a site rather than its live pages, see how to get historical website data. For the full single-page ranking on the 100 sites, see best web scraping APIs.
The one that finds every URL on your site within the caps you set, gets each page past the site's bot protection, and lists the pages it missed. Firecrawl, Scrapfly, Bright Data and Apify return content from one crawl job; String splits URL discovery and page fetch into two calls. Test each on a sample of your own site before you commit.
Start from the home page or a section URL, read robots.txt and the sitemap, follow links within the same host and path, and stop at a page cap. Then fetch each URL you found and keep a list of the ones that failed. A crawl API does the discovery and queueing for you; you still choose the scope and the caps.
Crawl it to get the URL list, then fetch every URL and extract what you need, as Markdown for reading or with a schema for structured fields. Scope the crawl to the section you need, cap pages and spend, honor any crawl delay, and check the failed list at the end so you know what your copy is missing.
Crawling finds pages: it follows links and sitemaps to build a list of URLs. Scraping takes content from pages: it fetches a URL and extracts text or fields. A full-site job does both, usually crawling first and scraping each result.
Both. The sitemap lists pages that no link reaches, and links find pages the sitemap leaves out. In our check of 100 bot-protected sites, 23 declared no sitemap in robots.txt, and 38 of the 77 that did would not serve the sitemap file to a plain HTTP client, so a crawler that relies only on the sitemap can miss the whole site.
It depends on the billing unit. Firecrawl charges 1 credit per page, Bright Data $1.50 per 1,000 requests pay as you go, and Apify estimates $0.20 to $5 per 1,000 pages depending on whether it renders a browser. String quotes each crawl before it runs; our 10-page test crawl was billed $0.003, and fetching content costs $0.20 to $0.30 per 1,000 standard requests.
Compare the crawl's URL count with the site's sitemap count for the same section, read the status of every URL, and check that pages returning 200 hold real content. A job marked complete has only finished its queue. Keep the list of URLs that failed and retry them.
It depends on the site's terms, the kind of data, copyright and where you operate, and robots.txt tells you what the owner allows. Public pages carry less risk than pages behind a login. See is web scraping legal and does robots.txt legally prevent scraping for the detail.