NewLaunching String Web Access APIRead the manifesto →
← Answers

Build vs buy web scraping infrastructure

String team · Updated September 28, 2026

Build web scraping infrastructure when your targets are few, unprotected and stable, and scraping is part of what you sell. Buy when bot-protected sites are in scope, because that is where an in-house stack spends most of its money. The fetch code is cheap. The cost sits in proxies, browsers, solvers, and the engineer who repairs the pipeline every time a target changes its defences. In String's Web Data Frontier Benchmark (September 16, 2026 run), 18 of the 100 bot-protected sites defeated most of the 16 scraping APIs tested, and those vendors do unblocking full time.

The decision rule

Count your targets and check what protects them. If every target returns its content to a plain HTTP client from a cloud server, build: an open-source framework and a small server will do, and no vendor adds much. If even a few targets sit behind Cloudflare, DataDome, Akamai, PerimeterX or Kasada, price the build honestly with the model below before you commit, because the protected sites set the cost of the whole system.

Scrapy, Playwright and Crawl4AI are free and good. None of them ships a proxy pool or a CAPTCHA solver. You buy or build those parts separately, and you own the result when they stop working.

A cost model you can fill in

Put your own numbers in each row and compare the monthly total against a vendor quote for the same volume and the same success rate.

Cost line What to count Build Buy (API)
Engineering to build Weeks to first production data, times loaded engineer cost You pay Integration only
Proxy bandwidth GB per month, split datacenter vs residential You pay In the per-request rate
Browser compute Pages that need JavaScript rendering, times CPU-seconds per page You pay In the per-request rate
Maintenance Engineer hours per month fixing selectors, fingerprints and blocks You pay Vendor fixes access; you keep parsers
Failures and retries Retries per successful page, times the cost of each attempt You pay every attempt Depends on vendor billing
Monitoring Alerts and checks that tell you a site broke before the data team does You build it Partly included

Two rows decide most outcomes. Maintenance is a standing cost, and it grows with the number of protected targets. Failure cost is invisible until you measure it: a site that passes three loads in five costs you five attempts for three pages, and each attempt burns bandwidth and compute.

For the buy column, use published rates. String charges per 1,000 successful requests: from $0.20 to $6.00 depending on plan, fetch path (plain HTTP or a rendered browser) and proxy class, on plans that start at $20 a month, with 5,000 standard requests free (pricing). Most failed requests are not billed.

Why the hard tail decides the answer

The benchmark sent 5 loads to each of 100 bot-protected sites through 16 APIs on September 16, 2026. A load passed only when the response contained a marker from the real page, so a CAPTCHA page with a 200 status counted as a failure.

  • On 18 of the 100 sites, more than half of the 16 APIs failed most of their loads.
  • On 19 sites, the average success rate across all 16 APIs was under 50%.
  • No site passed every load on all 16 APIs.
  • Temu defeated all 16, String included. String passed every load on 13 of the other 17 hard sites and failed Idealista.

The spread by anti-bot system is wide. On the 15 Cloudflare sites, results ran from 8% to 100%. On the 19 DataDome sites, from 4.2% to 92.6%. On the 20 Akamai sites, from 28% to 98%. Those are specialist vendors. An in-house stack starts below them and has to catch up with the same fingerprinting and challenge work.

Practitioners say the same. A comment in r/webscraping put it this way: small scrapers for unprotected sites "are mostly gone or worth a lot less", but once scale or bot mitigation grows, "LLMs can't one or multiple-shot anything." AI coding tools made the easy part of scraping cheap. The hard part still costs what it did.

When building wins

  • Fewer than a handful of targets, and none is bot-protected.
  • The targets rarely change layout, and a day of stale data costs little.
  • Scraping is your product or your moat, so owning the stack is the point.
  • Very high, steady volume on a fixed set of easy sites, where a per-request rate adds up.
  • A requirement forces you to own the network path end to end.

When buying wins

  • Any meaningful share of targets sits behind a commercial anti-bot system.
  • The target list grows or changes each quarter.
  • Nobody on the team wants to be on call for a broken proxy pool.
  • The data feeds a product, a model or an agent, and collection is plumbing.
  • You need a success rate you can measure and put in a contract.

A common setup is hybrid: buy access (proxies, rendering, unblocking) through an API such as the Web Access API, and keep parsing and business logic in-house. If you want neither, a managed service delivers the dataset and owns the pipeline. Our page on managed scraping services vs scraping APIs covers that split.

What to ask a vendor

  • How do you measure success: on page content or on HTTP status?
  • What is your success rate on my targets? Run a trial on your own URL list, not the vendor's demo sites.
  • Do I pay for blocked, empty or timed-out requests?
  • What does the SLA cover, and what happens when you miss it?
  • Where do you send my requests, what do you log, and how long do you keep it?
  • What happens to my data and my code if I leave?

FAQ

Is it cheaper to build or buy web scraping infrastructure?

Building is cheaper when your targets are few, stable and unprotected. Buying is cheaper once bot-protected sites are in scope, because proxies, browser compute, CAPTCHA handling and maintenance make up most of an in-house stack's cost. In String's September 16, 2026 benchmark, 18 of 100 protected sites defeated most of the 16 APIs tested.

How should an engineering team decide between building scraping infrastructure and buying an API?

List the targets and check which anti-bot system protects each one. Fill in a monthly cost model for engineering, proxy bandwidth, browser compute, maintenance and failed attempts, then compare it with a vendor's rate at the same success rate on your own URLs. Build if the targets are easy and scraping is your product; buy access if protected sites are in scope.

Apify vs building my own scraper: which is cheaper long term?

Run the same cost model for both. A self-built scraper costs engineering, proxies, compute and maintenance; a platform charges for usage and still leaves you to maintain some scraper logic. In the September 16, 2026 benchmark Apify passed 77.4% of 500 loads on 100 protected sites, so measure the retries and failures on your own targets before you compare totals.

Are open-source scraping frameworks enough?

Scrapy, Playwright and Crawl4AI are free and cover fetching, crawling and browser control. None of them ships a proxy pool or a CAPTCHA solver, so protected sites still need those parts, bought or built. On unprotected sites, an open-source framework is often all you need.

What should I ask a scraping vendor before I buy?

Ask whether success is measured on page content or HTTP status, and test the vendor on your own URLs. Ask whether you pay for blocked or empty requests, what the SLA covers and what the remedy is, and how the vendor handles and retains your request data.

Sources

  • Web Data Frontier Benchmark, run of September 16, 2026: 16 APIs, 100 bot-protected sites, 5 loads per site, 90-second timeout. Results and method at usestring.ai/benchmark; harness and raw results on GitHub.
  • String rates from usestring.ai/pricing, checked September 28, 2026. Billing on failed requests from the String docs.
  • r/webscraping comment, permalink.
Get your API key →Explore the Web Access API
© 2026 StringEU and UK GDPR Article 27 representative — appointment verified by EuverifyBuilt in New York City 🗽 🍎