Build web scraping infrastructure when your targets are few, unprotected and stable, and scraping is part of what you sell. Buy when bot-protected sites are in scope, because that is where an in-house stack spends most of its money. The fetch code is cheap. The cost sits in proxies, browsers, solvers, and the engineer who repairs the pipeline every time a target changes its defences. In String's Web Data Frontier Benchmark (September 16, 2026 run), 18 of the 100 bot-protected sites defeated most of the 16 scraping APIs tested, and those vendors do unblocking full time.
Count your targets and check what protects them. If every target returns its content to a plain HTTP client from a cloud server, build: an open-source framework and a small server will do, and no vendor adds much. If even a few targets sit behind Cloudflare, DataDome, Akamai, PerimeterX or Kasada, price the build honestly with the model below before you commit, because the protected sites set the cost of the whole system.
Scrapy, Playwright and Crawl4AI are free and good. None of them ships a proxy pool or a CAPTCHA solver. You buy or build those parts separately, and you own the result when they stop working.
Put your own numbers in each row and compare the monthly total against a vendor quote for the same volume and the same success rate.
| Cost line | What to count | Build | Buy (API) |
|---|---|---|---|
| Engineering to build | Weeks to first production data, times loaded engineer cost | You pay | Integration only |
| Proxy bandwidth | GB per month, split datacenter vs residential | You pay | In the per-request rate |
| Browser compute | Pages that need JavaScript rendering, times CPU-seconds per page | You pay | In the per-request rate |
| Maintenance | Engineer hours per month fixing selectors, fingerprints and blocks | You pay | Vendor fixes access; you keep parsers |
| Failures and retries | Retries per successful page, times the cost of each attempt | You pay every attempt | Depends on vendor billing |
| Monitoring | Alerts and checks that tell you a site broke before the data team does | You build it | Partly included |
Two rows decide most outcomes. Maintenance is a standing cost, and it grows with the number of protected targets. Failure cost is invisible until you measure it: a site that passes three loads in five costs you five attempts for three pages, and each attempt burns bandwidth and compute.
For the buy column, use published rates. String charges per 1,000 successful requests: from $0.20 to $6.00 depending on plan, fetch path (plain HTTP or a rendered browser) and proxy class, on plans that start at $20 a month, with 5,000 standard requests free (pricing). Most failed requests are not billed.
The benchmark sent 5 loads to each of 100 bot-protected sites through 16 APIs on September 16, 2026. A load passed only when the response contained a marker from the real page, so a CAPTCHA page with a 200 status counted as a failure.
The spread by anti-bot system is wide. On the 15 Cloudflare sites, results ran from 8% to 100%. On the 19 DataDome sites, from 4.2% to 92.6%. On the 20 Akamai sites, from 28% to 98%. Those are specialist vendors. An in-house stack starts below them and has to catch up with the same fingerprinting and challenge work.
Practitioners say the same. A comment in r/webscraping put it this way: small scrapers for unprotected sites "are mostly gone or worth a lot less", but once scale or bot mitigation grows, "LLMs can't one or multiple-shot anything." AI coding tools made the easy part of scraping cheap. The hard part still costs what it did.
A common setup is hybrid: buy access (proxies, rendering, unblocking) through an API such as the Web Access API, and keep parsing and business logic in-house. If you want neither, a managed service delivers the dataset and owns the pipeline. Our page on managed scraping services vs scraping APIs covers that split.
Building is cheaper when your targets are few, stable and unprotected. Buying is cheaper once bot-protected sites are in scope, because proxies, browser compute, CAPTCHA handling and maintenance make up most of an in-house stack's cost. In String's September 16, 2026 benchmark, 18 of 100 protected sites defeated most of the 16 APIs tested.
List the targets and check which anti-bot system protects each one. Fill in a monthly cost model for engineering, proxy bandwidth, browser compute, maintenance and failed attempts, then compare it with a vendor's rate at the same success rate on your own URLs. Build if the targets are easy and scraping is your product; buy access if protected sites are in scope.
Run the same cost model for both. A self-built scraper costs engineering, proxies, compute and maintenance; a platform charges for usage and still leaves you to maintain some scraper logic. In the September 16, 2026 benchmark Apify passed 77.4% of 500 loads on 100 protected sites, so measure the retries and failures on your own targets before you compare totals.
Scrapy, Playwright and Crawl4AI are free and cover fetching, crawling and browser control. None of them ships a proxy pool or a CAPTCHA solver, so protected sites still need those parts, bought or built. On unprotected sites, an open-source framework is often all you need.
Ask whether success is measured on page content or HTTP status, and test the vendor on your own URLs. Ask whether you pay for blocked or empty requests, what the SLA covers and what the remedy is, and how the vendor handles and retains your request data.