NewLaunching String Web Access APIRead the manifesto →
← Answers

How to crawl an entire website with an API

String team · Updated October 5, 2026

The best API for crawling an entire website is the one that does two jobs well on your site: it finds every URL, from the sitemap and from links, inside a page cap and a spend cap you set; and it gets each page's content past the site's bot protection, then tells you which pages it missed. Firecrawl's /crawl, Scrapfly's Crawler API, Bright Data's Crawl API and Apify's Website Content Crawler do both in one job. String runs them as two calls: a /sitemap job lists the URLs, then /fetch returns each page as Markdown. Test on your own site first. When we checked the 100 sites in String's benchmark on October 3, 2026, 38 of the 77 sites that publish a sitemap did not serve the sitemap file itself to a plain HTTP client.

A full-site crawl is two jobs

People who search for this want one of three things: a list of every URL on a site, the content of every page, or a vendor. The first two are separate jobs with separate failure modes, and it helps to keep them apart even when one API does both.

Job What it does How it fails
Discovery Reads robots.txt and the sitemap files, follows links from page to page, stays inside the scope you set Misses pages that no sitemap lists and no link reaches; follows links into calendars and filters forever; gets blocked on the sitemap file
Content Fetches each discovered URL and returns HTML, Markdown or JSON Gets a block page or an empty 200 instead of the content; hits a rate limit halfway through; renders a page before its JavaScript has loaded

Discovery costs little per URL. Content is a long series of single-page fetches, so a crawl is only as good as the API's success on one protected page, repeated thousands of times.

What the sitemap files on 100 protected sites showed

String's benchmark tests single-page fetches on 100 bot-protected sites: retail, travel, real estate, jobs, tickets, social. It does not test crawl jobs. On October 1 and 3, 2026 we went back to those 100 sites and checked the first thing a crawler reads: robots.txt and the sitemap it points to.

  • 77 of 100 sites declare at least one sitemap in robots.txt, 1,334 sitemap lines between them. The other 23, including Amazon, LinkedIn, Reddit, Yelp, Indeed and Ticketmaster, declare none. On those, discovery has to follow links.
  • 16 of 100 set a Crawl-delay, from 0.2 to 15 seconds. A crawler that honors it on a 10,000-page site at 10 seconds a page needs more than a day. Our guide to handling rate limits when scraping covers pacing.
  • On 38 of the 77, a plain HTTP client with a Chrome User-Agent did not get the sitemap. 31 answered HTTP 403, one answered 418, one redirected away, and five returned a 200 with no sitemap in it. Two more had declared sitemaps that no longer exist (404).
  • String's fetch returned a valid sitemap for 29 of those 38. The same bot protection that guards the product pages often guards the sitemap file, so a crawler that cannot get past it may never see the URL list.

The site list is in the public benchmark harness, in src/tests.const.ts. For each site we fetched the first sitemap that robots.txt declares.

Developers report the same problems in public issue threads. One Firecrawl user reported a crawl where the status endpoint kept returning scraping "but the completed value never changes" (firecrawl #1679). A Crawl4AI user running the docs' own example wrote: "I cannot scrape more than 2 pages - it gets stuck" (crawl4ai #1017). And on r/webscraping, someone copying a 100-page site said its scripts "mess up the normal wget and httrack downloading apps" (thread).

Crawl endpoints, from each vendor's docs

Checked against each vendor's documentation on October 3, 2026. String's benchmark measures single-page fetches, not crawl jobs, so this table compares documented features and does not rank anyone.

API Crawl endpoint Returns page content Page cap Depth control Billing Results kept
String POST /v1/sitemap, then POST /v1/fetch per URL URLs from /sitemap; content from /fetch maxPages 1 to 10,000 maxDepth 1 to 100 Quote first; billed per page crawled at your plan's rate; fetch billed per successful request URL lists stay available after the job
Firecrawl POST /v2/crawl Yes, Markdown or JSON limit, default 10,000 maxDiscoveryDepth 1 credit per page; JSON mode 4 more credits per page 24 hours
Scrapfly Crawler API Yes, several formats plus WARC and HAR page_limit (0 = no limit within plan) max_depth Sum of the Web Scraping API calls the crawl makes; max_api_credit cap Not stated
Bright Data Crawl API Yes, Markdown, text, HTML or JSON Not documented publicly Not documented publicly $1.50 per 1,000 requests pay as you go Not stated
Apify Website Content Crawler Actor run Yes maxCrawlPages Max crawling depth Compute units, about $0.20 per 1,000 pages over raw HTTP and $0.50 to $5 with a browser, by Apify's estimate Dataset
Zyte API None; /extract takes one URL Per URL You bring the crawler You bring the crawler Per request Not applicable

Sources: String sitemap reference, Firecrawl crawl docs, Scrapfly Crawler API, Bright Data Crawl API pricing, Apify Website Content Crawler, Zyte API reference. Bright Data's product page prices the same tiers per 1,000 "records" while its pricing page says "requests"; we quote the pricing page.

If you would rather run the crawler yourself, Scrapy, Crawlee and Crawl4AI handle discovery and queueing for free, and you can point their page fetches at any scraping API.

What to check before you pick one

  1. Does it return content, or only URLs? Both are useful. A URL list lets you diff a site over time and fetch only what changed.
  2. Can you set a page cap and a spend cap? Without them, a calendar or a faceted search can turn a 2,000-page site into 200,000 URLs. Firecrawl, for example, checks your balance against its default limit of 10,000 pages and returns a 402 before the crawl starts if the balance is short.
  3. Can you scope it? Path prefixes, include and exclude patterns, subdomains. Most crawls should stay under one path such as /docs or /products.
  4. Does it read the sitemap and follow links? You need both: the sitemap finds pages no link reaches, and links find pages the sitemap forgot. On the 23 benchmark sites without a declared sitemap, links are all you have.
  5. Does it get past the bot protection on your site? The content step is thousands of single-page fetches. Check a sample of real pages on your target, not the vendor's demo site. How to evaluate a web scraping API explains how to size that test.
  6. Does it list the pages it missed? Firecrawl's own docs warn that "a crawl that finishes is not the same as a crawl that reached every page." Ask for a per-URL status and an error list, not a single "completed" flag.
  7. Are failed pages billed? Pricing differs on this more than on the headline rate.
  8. How long are results kept? Firecrawl keeps crawl results for 24 hours after completion; download them before they expire.

How to crawl a site with String, step by step

  1. Ask for a quote. POST /v1/sitemap with the start URL, maxPages, maxDepth, an optional pathPrefix, and useSitemap: true to seed from /sitemap.xml as well as links. The crawl stays on the start URL's hostname. Nothing is crawled or charged yet.
  2. Read the quote. It returns estimatedPages and estimatedCostUsd, a ceiling for the job. For a 10-page crawl of our own site the quote was $0.053; the job was billed $0.003.
  3. Approve. POST /v1/sitemap/{jobId}/approve starts the crawl and holds the quote against your balance. budgetUsd caps spend; leave it at the default, which is the quote. When we set it far below the quote ($0.01 against $0.053), the job stopped with token_cap_exceeded after 2 of 10 pages, so set it at or above the quote.
  4. Poll the status until it is completed, token_cap_exceeded, canceled or failed.
  5. Page through the URLs. GET /v1/sitemap/{jobId}/urls returns each URL with its statusCode, depth, parentUrl and sourceType (seed, sitemap or HTML link), up to 5,000 per page.
  6. Fetch the content. POST /v1/fetch with format: "markdown" for each URL that answered 200. Keep under the account rate of 60 requests per second.
  7. Keep the missing list. Every URL with a status other than 200 is a page your copy does not have. Retry those later, or log them.

Python

We ran this on October 3, 2026 against https://usestring.ai/answers with a 10-page cap: the crawl completed, was billed $0.003, and all 9 pages fetched as Markdown returned 200. Without a pathPrefix the crawl left /answers and covered the rest of the site, which is the default behavior: it stays on the hostname, not the path.

"""Crawl a whole site in two steps: discover every URL, then fetch each page as Markdown.

    pip install requests
    export STRING_API_KEY=...
    python crawl_site.py https://example.com/docs 500
"""
import json
import os
import sys
import time

import requests

API = "https://request.usestring.ai/v1"
HEADERS = {"Authorization": f"Bearer {os.environ['STRING_API_KEY']}"}
MAX_QUOTE_USD = 5.00  # refuse to start a crawl quoted above this


def discover(start_url: str, max_pages: int) -> list:
    quote = requests.post(f"{API}/sitemap", headers=HEADERS, timeout=60, json={
        "url": start_url, "maxPages": max_pages, "maxDepth": 5, "useSitemap": True,
    }).json()
    print(f"quote: up to {quote['estimatedPages']} pages, at most ${quote['estimatedCostUsd']}")
    if float(quote["estimatedCostUsd"]) > MAX_QUOTE_USD:
        sys.exit("quote is above MAX_QUOTE_USD; lower maxPages or raise the cap")
    job = quote["jobId"]
    requests.post(f"{API}/sitemap/{job}/approve", headers=HEADERS, timeout=60).raise_for_status()
    while True:
        status = requests.get(f"{API}/sitemap/{job}", headers=HEADERS, timeout=60).json()
        if status["status"] not in ("awaiting_approval", "running"):
            break
        time.sleep(5)
    print(f"crawl {status['status']}, cost ${status.get('costUsd')}")
    urls, offset = [], 0
    while True:
        page = requests.get(f"{API}/sitemap/{job}/urls", headers=HEADERS, timeout=60,
                            params={"limit": 1000, "offset": offset}).json()
        urls += page["urls"]
        offset += len(page["urls"])
        if not page["urls"] or offset >= page["total"]:
            return urls


def fetch_markdown(url: str) -> dict:
    r = requests.post(f"{API}/fetch", headers=HEADERS, timeout=120, json={"url": url, "format": "markdown"})
    if r.status_code == 429:  # over the account rate: wait and try once more
        time.sleep(2)
        r = requests.post(f"{API}/fetch", headers=HEADERS, timeout=120, json={"url": url, "format": "markdown"})
    # Markdown comes back as the response body; the site's own status is in X-Status-Code.
    return {"url": url, "status": int(r.headers.get("X-Status-Code", r.status_code)), "markdown": r.text}


if __name__ == "__main__":
    start, limit = sys.argv[1], int(sys.argv[2])
    found = discover(start, limit)
    pages = [u["url"] for u in found if u["statusCode"] == 200 and not u["isSitemap"]]
    failed = [u for u in found if u["statusCode"] != 200 and not u["isSitemap"]]
    print(f"discovered {len(found)} URLs: {len(pages)} answered 200, {len(failed)} did not")
    with open("pages.jsonl", "w") as out:
        for url in pages:
            result = fetch_markdown(url)
            out.write(json.dumps(result) + "\n")
            print(result["status"], len(result["markdown"]), url)
    for u in failed:  # keep this list: these pages are missing from your copy
        print("MISSING", u["statusCode"], u["url"], "from", u["parentUrl"])

TypeScript

The same flow for Node 18+ or Bun. Same run, same result on October 3, 2026.

// Crawl a whole site in two steps: discover every URL, then fetch each page as Markdown.
//   export STRING_API_KEY=...
//   npx tsx crawl_site.ts https://example.com/docs 500   (or: bun crawl_site.ts ...)
import { writeFileSync } from "node:fs";

const API = "https://request.usestring.ai/v1";
const HEADERS = { Authorization: `Bearer ${process.env.STRING_API_KEY}`, "Content-Type": "application/json" };
const MAX_QUOTE_USD = 5; // refuse to start a crawl quoted above this
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));

type Found = { url: string; statusCode: number; isSitemap: boolean; parentUrl: string | null };

async function discover(startUrl: string, maxPages: number): Promise<Found[]> {
  const quote = await fetch(`${API}/sitemap`, {
    method: "POST",
    headers: HEADERS,
    body: JSON.stringify({ url: startUrl, maxPages, maxDepth: 5, useSitemap: true }),
  }).then((r) => r.json());
  console.log(`quote: up to ${quote.estimatedPages} pages, at most $${quote.estimatedCostUsd}`);
  if (Number(quote.estimatedCostUsd) > MAX_QUOTE_USD) throw new Error("quote is above MAX_QUOTE_USD");
  await fetch(`${API}/sitemap/${quote.jobId}/approve`, { method: "POST", headers: HEADERS });
  let status;
  do {
    await sleep(5000);
    status = await fetch(`${API}/sitemap/${quote.jobId}`, { headers: HEADERS }).then((r) => r.json());
  } while (["awaiting_approval", "running"].includes(status.status));
  console.log(`crawl ${status.status}, cost $${status.costUsd}`);
  const urls: Found[] = [];
  for (let offset = 0; ; ) {
    const page = await fetch(`${API}/sitemap/${quote.jobId}/urls?limit=1000&offset=${offset}`, { headers: HEADERS }).then((r) => r.json());
    urls.push(...page.urls);
    offset += page.urls.length;
    if (!page.urls.length || offset >= page.total) return urls;
  }
}

async function fetchMarkdown(url: string) {
  const res = await fetch(`${API}/fetch`, { method: "POST", headers: HEADERS, body: JSON.stringify({ url, format: "markdown" }) });
  // Markdown comes back as the response body; the site's own status is in X-Status-Code.
  return { url, status: Number(res.headers.get("x-status-code") ?? res.status), markdown: await res.text() };
}

const [start, limit] = [process.argv[2], Number(process.argv[3])];
const found = await discover(start, limit);
const pages = found.filter((u) => u.statusCode === 200 && !u.isSitemap).map((u) => u.url);
const failed = found.filter((u) => u.statusCode !== 200 && !u.isSitemap);
console.log(`discovered ${found.length} URLs: ${pages.length} answered 200, ${failed.length} did not`);
const results = [];
for (const url of pages) {
  const r = await fetchMarkdown(url);
  results.push(r);
  console.log(r.status, r.markdown.length, url);
}
writeFileSync("pages.jsonl", results.map((r) => JSON.stringify(r)).join("\n") + "\n");
for (const u of failed) console.log("MISSING", u.statusCode, u.url, "from", u.parentUrl); // pages missing from your copy

The failure mode to watch for: a crawl that says it finished

A crawl job that reports completed has finished its queue. It has not proved it reached every page. Three checks catch most gaps:

  • Compare against the sitemap. Count the URLs in the site's sitemap files and the URLs your crawl found under the same path. A crawl far below the sitemap count stopped early or was blocked.
  • Read the per-URL statuses. A 403, a 429 or a status of 0 means the page is missing from your copy. String's /sitemap results carry a status for every URL; keep the list and retry it.
  • Check the content, not just the status. A block page can come back as HTTP 200. How to tell a bot block from a real page lists the checks.

Where String fits

String's Web Access API suits teams that want the URL list and the content as separate, inspectable steps: discovery with a quote before anything is charged, then content from the fetch API that passed 485 of 500 requests (97.0%) on the 100 bot-protected sites in our September 16, 2026 benchmark, the highest of 16 providers. We build String and run that benchmark, so test it on your own site.

  • Limits: up to 10,000 pages and depth 100 per crawl job, same hostname; fetch at 60 requests per second per organization with no cap on parallel requests.
  • Price: crawl pages are billed at your plan's rate after a quote; fetch is $0.30 per 1,000 successful standard requests on the $20 Starter plan and $0.20 on the $100 Growth plan, and you pay only when String returns content. The first 5,000 standard requests are free (pricing).
  • Integrations: REST from any language, the MCP server for agents (it exposes the crawl job as a tool), and Apify.
  • When one call is simpler: if you want Markdown for a small docs site in a single job and do not need the URL list, an API that returns content inside the crawl does it in fewer steps.

To pull an archived copy of a site rather than its live pages, see how to get historical website data. For the full single-page ranking on the 100 sites, see best web scraping APIs.

FAQ

What is the best API for crawling an entire website?

The one that finds every URL on your site within the caps you set, gets each page past the site's bot protection, and lists the pages it missed. Firecrawl, Scrapfly, Bright Data and Apify return content from one crawl job; String splits URL discovery and page fetch into two calls. Test each on a sample of your own site before you commit.

How do I crawl all pages of a website?

Start from the home page or a section URL, read robots.txt and the sitemap, follow links within the same host and path, and stop at a page cap. Then fetch each URL you found and keep a list of the ones that failed. A crawl API does the discovery and queueing for you; you still choose the scope and the caps.

How do I scrape a whole website?

Crawl it to get the URL list, then fetch every URL and extract what you need, as Markdown for reading or with a schema for structured fields. Scope the crawl to the section you need, cap pages and spend, honor any crawl delay, and check the failed list at the end so you know what your copy is missing.

What is the difference between crawling and scraping?

Crawling finds pages: it follows links and sitemaps to build a list of URLs. Scraping takes content from pages: it fetches a URL and extracts text or fields. A full-site job does both, usually crawling first and scraping each result.

Should I use the sitemap or follow links?

Both. The sitemap lists pages that no link reaches, and links find pages the sitemap leaves out. In our check of 100 bot-protected sites, 23 declared no sitemap in robots.txt, and 38 of the 77 that did would not serve the sitemap file to a plain HTTP client, so a crawler that relies only on the sitemap can miss the whole site.

How much does it cost to crawl a website with an API?

It depends on the billing unit. Firecrawl charges 1 credit per page, Bright Data $1.50 per 1,000 requests pay as you go, and Apify estimates $0.20 to $5 per 1,000 pages depending on whether it renders a browser. String quotes each crawl before it runs; our 10-page test crawl was billed $0.003, and fetching content costs $0.20 to $0.30 per 1,000 standard requests.

How do I know my crawl got every page?

Compare the crawl's URL count with the site's sitemap count for the same section, read the status of every URL, and check that pages returning 200 hold real content. A job marked complete has only finished its queue. Keep the list of URLs that failed and retry them.

Is it legal to crawl an entire website?

It depends on the site's terms, the kind of data, copyright and where you operate, and robots.txt tells you what the owner allows. Public pages carry less risk than pages behind a login. See is web scraping legal and does robots.txt legally prevent scraping for the detail.

Get your API key →Explore the Web Access API
© 2026 StringEU and UK GDPR Article 27 representative — appointment verified by EuverifyBuilt in New York City 🗽 🍎