NewLaunching String Web Access APIRead the manifesto →
← Answers

How to evaluate a web scraping API before you buy

String team · Updated September 30, 2026

A web scraping API fetches a web page for you through one HTTP call and handles the proxies, browser rendering and anti-bot challenges behind that call. Evaluate one with a trial on your own URLs, not the vendor's demo pages. Use enough of the pages you need, fetch each one several times, and count a request as a pass only when the body contains content that belongs on that page. Then compare cost per passed page, latency on passed pages, and what the vendor does when a site changes. Small trials mislead. In String's Web Data Frontier Benchmark run of September 16, 2026, a random 10-site trial put two providers within 5 points of each other in the wrong order 40% of the time.

What a web scraping API does

You send a URL. The API picks an IP address, often a residential or mobile one for protected sites, and sends a request that looks like a normal browser. It renders JavaScript when the page needs it and gets past bot challenges from systems such as Cloudflare, DataDome, PerimeterX, Akamai and Kasada. It returns the page as HTML, Markdown or extracted JSON. You pay for the calls, not for the proxy network and the browser fleet behind them.

That is also why evaluation is hard. The work the API does is invisible, and it differs by site. An API that passes every retail site in your list can fail every travel site. A buyer comparing e-commerce APIs on r/WebScrapingInsider put it this way: "I don't want to save money on Amazon only to discover that the same provider is expensive or unreliable on Walmart and the smaller sites." Another team on r/webscraping kept three APIs at once, partly "to be able to have a variety of web scrapers for different sites."

Why a small trial gives the wrong answer

The benchmark sends every provider the same 100 bot-protected pages, five attempts each, and counts a pass only when the status is 2xx and the body contains a text string from that page. We used the September 16, 2026 results to test what a smaller trial would have told a buyer.

  • Close providers swap places. In the run, 23 pairs of providers finished within 5 points of each other. We drew 4,000 random samples of target sites and compared each pair on each sample. With 10 sites, the sample put the pair in the wrong order 39.5% of the time and tied them another 5.5%. With 20 sites it was wrong 36.3% of the time, and with 50 sites 26.7%.
  • One provider's score moves a lot. A provider that passed 77.4% of all 500 requests scored anywhere from 58% to 96% on a random 10-site sample (the 5th to 95th percentile of 10,000 draws).
  • One attempt per URL is not enough. In 299 of the 1,600 provider and site cells (18.7%), the five attempts did not agree: some passed and some failed. A trial that fetches each URL once calls those sites pass or fail by chance.

So a gap of a few points on a small trial is noise. Either test more URLs or treat the close providers as tied and decide on cost, latency and support.

How to run the trial

  1. Pick the URLs from your own production list. Include the hardest sites, not only the easy ones, and weight the list toward where the value is. Aim for 50 or more pages if you need to separate providers that are close.
  2. Write a pass rule for each URL. A 2xx status plus a string or selector that only the real page carries, such as the product title or a price field. A status code alone is not a pass: in the same run, 7.7% of 2xx responses did not contain the page's content. Our guide to checking that scraped data is complete and accurate covers the checks.
  3. Fetch each URL several times, over more than one day. Five attempts per URL separate a site that is blocked from one that fails now and then.
  4. Send the same URLs to every candidate in the same window. Sites change their protection. A trial of vendor A this week against vendor B last month compares two different webs.
  5. Record every attempt: status, pass or fail, latency, bytes and the billed cost. Take the billed cost from the provider's usage report, or from a response header where it gives one.
  6. Score per site, then overall. Success rate, cost per passed page, and median and 95th-percentile latency on passed pages. Look at the per-site table before the average: one site you depend on can matter more than ten you do not.

The benchmark harness is open source. Each target is a fixture of name, url and containsText in src/tests.const.ts, and each provider is a small adapter in src/providers/. Replace the fixtures with your own URLs and markers, add the API keys you want to test to .env, and run it. It applies the pass rule above and reports success rate and latency per provider and per site.

What to check beyond the trial

These are the questions enterprise data teams raise when they replace an in-house scraping stack:

  • Billing unit. What counts as a billable request, whether a blocked or failed request is billed, and which features cost more (browser rendering, premium proxies). Your trial's cost per passed page answers most of this with real numbers.
  • Limits at your volume. Requests per second, concurrent requests, and what happens when you exceed them.
  • What breaks when a site changes. An API returns the page. The parser that turns it into fields is still yours, unless you use the vendor's structured extraction or a managed service with a repair commitment. Managed scraping service vs scraping API compares the two.
  • Output. HTML, Markdown, JSON against a schema, or screenshots, and whether you can switch per request.
  • Data handling. Whether the vendor stores your requests or responses, for how long, and which security reports it can share. Ask for documents, not a sales answer.
  • Migration. Run the API beside your current stack on the same URLs for two weeks or more before you switch traffic, and keep your parsing code separate from the fetch call so you can change vendors later.

If the question is whether to replace the in-house stack at all, build vs buy web scraping infrastructure covers that decision.

Where String fits

String's Web Access API passed 97.0% of requests in the September 16, 2026 run, the highest of the 16 providers tested. We run that benchmark ourselves, so treat it as one input and test on your own URLs. The first 5,000 standard requests are free, and you only pay when String returns content. Current rates are on the pricing page, and our security and compliance status is on the trust page.

The benchmark measures bot-protected pages. If your sites have no bot protection, most APIs will pass most requests, and price, latency and output format should decide. For the full ranking on the 100 sites, see best web scraping APIs. For why the pass rule decides what a success rate means, see the web scraping benchmark problem.

FAQ

What is a web scraping API?

A web scraping API is a service that fetches a web page for you through one HTTP call. It handles proxy rotation, JavaScript rendering and anti-bot challenges, and returns the page as HTML, Markdown or structured JSON. You pay per request instead of running proxies and browsers yourself.

How should I evaluate a web scraping API?

Run a trial on your own URLs, not the vendor's demo pages. Fetch each URL several times, count a pass only when the body contains content from that page, and compare success rate, cost per passed page and latency, site by site. Then check billing rules, rate limits, output formats and what happens when a site changes.

How many URLs should a scraping API trial include?

Enough to separate the providers you are choosing between. In our September 16, 2026 benchmark run, a random 10-site trial ordered providers within 5 points of each other wrongly 39.5% of the time, and a 50-site trial 26.7% of the time. With a small list, treat close results as a tie.

How do I measure a provider's real success rate before buying?

Send it your own URLs, five times each, over more than one day, and apply your own pass rule: a 2xx status plus a page-specific string in the body. Divide passed requests by all requests. Do not rely on the provider's dashboard, because it may count any 2xx response as a success.

Are vendor-published scraping benchmarks reliable?

Treat them as one input. Check the pass rule, the target list and whether the harness is public, because those decide what the number means. A benchmark on unprotected pages says little about protected ones. String publishes its harness and raw results so anyone can rerun them, and it is still a vendor-run benchmark.

What should an enterprise data team evaluate when replacing an in-house scraping stack?

Success rate and cost per passed page on your own URLs, limits at your volume, what breaks when a site changes and who fixes it, output formats, data handling and security documents, and a migration plan. Run the new service beside the old stack on the same URLs for two weeks or more before you move traffic.

Is there a free trial for web scraping APIs?

Most providers offer free credits or a free tier. Check whether failed or blocked requests use up those credits, and whether rendering or premium proxies cost several credits per request, because both change how far a trial goes. String gives 5,000 standard requests free and bills only when it returns content.

How long should a web scraping API trial run?

At least several days. Bot protection can react to traffic patterns over time, so a result from one session may not hold the next day. Fetch each URL on more than one day, and send the same URLs to every candidate in the same window so the comparison is fair.

Sources

Get your API key →Explore the Web Access API
© 2026 StringEU and UK GDPR Article 27 representative — appointment verified by EuverifyBuilt in New York City 🗽 🍎