Check scraped data at three levels, because each level catches failures the other two miss. Check each response: did the page you asked for come back, or a block page with a 200 status? Check each record: are the required fields present, typed and in range? Check each run: did row counts, fill rates and duplicate counts stay close to the last run that passed? Then compare a random sample of records with the live pages by hand, because a value can be wrong and still look right. The status code alone proves nothing. In String's Web Data Frontier Benchmark run of September 16, 2026, 443 of the 5,769 responses that came back with a 2xx status (7.7%) did not contain the content the page should hold.
The benchmark sends 500 requests per provider to 100 target pages, five attempts per page, for 16 scraping providers. An attempt passes only when the status is 2xx and the body contains a text string that belongs on that page, such as the product title. In the September 16, 2026 run, 2,674 of 8,000 attempts failed. One failure in six (443) was a 2xx response without the expected content. A pipeline that trusts the status code would have stored all 443 as data.
The problem was not rare or specific to one vendor. All 16 providers returned at least 5 of these responses, and 54 of the 100 pages drew at least one. On a Temu product page, no attempt from any provider passed, and 50 of the 80 attempts returned a 2xx without the product title. On a Zara page, 46 attempts passed and 33 returned a 2xx without the expected text, on the same URL in the same run. So one good fetch says little about the next one. Check every response, not a sample.
Two limits apply to that number. It is a floor: a short marker can match a challenge page, as our guide to block pages that return HTTP 200 shows for Reddit. And the check cannot say why the content was missing. A block page and a changed listing both fail it.
Practitioners describe the same failure. One r/webscraping user who runs about 30 scrapers wrote that what annoys them most is scrapers that don't crash but "silently return bad data because a selector changed," and added: "Sometimes I don't notice for days." Another, who collects product prices across regions, described a source API that renamed a parameter, after which "all prices in one country had the wrong currency." Their per-brand, per-region price counts did not catch it. A field-level currency check would have.
Response. Before you parse anything, confirm the page is the page. Test for a marker that only the real page carries, count the items you expected, and compare the final URL with the one you requested. The HTTP 200 block page guide has the full checklist.
Record. Validate every row against a schema: required fields present, types correct, formats valid (dates, currencies, IDs). Add range checks (a price above zero and below a sane ceiling) and cross-field checks (a sale price is not above the list price). Store the source URL, the collection timestamp and the raw page with each record. When a number looks wrong later, you can open the exact page it came from.
Run. Compare each run with the last run that passed, per source. Watch the row count, the fill rate of each important field, the count of distinct keys and the duplicate count. Also watch for the opposite failure. A crawler that serves a cached or stale page returns a run that looks perfect: the same count, the same values. If a source that changes every day shows zero changed values, treat the run as suspect.
Coverage. Completeness needs a denominator from outside your scraper: the site's own "N results" figure, its sitemap, or a second source. Then ask where the gaps fall. A dataset that misses a quarter of locations is usable if the misses are random. The same dataset misleads you if every miss is in one segment, such as large chains.
Accuracy sample. Pick records at random from every source on every run and compare them with the live page by hand. Automated checks catch change. Only a person looking at the source catches a value that is wrong but plausible, such as a regular price stored as the sale price.
This script runs the record and run checks above on a list of dictionaries. It uses only the standard library. We ran it on Python 3.9 on September 30, 2026, and the output is below it.
import hashlib
REQUIRED = {"url": str, "title": str, "price": float, "currency": str}
CURRENCIES = {"USD", "EUR", "GBP"}
def check_record(r):
"""Record level: return a list of problems, empty if the record is good."""
problems = []
for field, kind in REQUIRED.items():
if not isinstance(r.get(field), kind) or r.get(field) in ("", None):
problems.append(f"{field} missing or wrong type")
if isinstance(r.get("price"), float) and not 0 < r["price"] < 100_000:
problems.append("price out of range")
if r.get("currency") not in CURRENCIES:
problems.append("unexpected currency")
if r.get("sale_price") and r["sale_price"] > r.get("price", 0):
problems.append("sale price above list price")
return problems
def check_run(records, last_good):
"""Run level: compare this run with the last run that passed."""
alerts = []
good = [r for r in records if not check_record(r)]
if len(good) < 0.8 * last_good["rows"]:
alerts.append(f"rows fell to {len(good)} from {last_good['rows']}")
fill = sum(1 for r in records if r.get("price")) / max(len(records), 1)
if fill < last_good["price_fill"] - 0.05:
alerts.append(f"price fill rate fell to {fill:.0%}")
keys = [r.get("url") for r in records]
if len(keys) != len(set(keys)):
alerts.append(f"duplicate keys: {len(keys) - len(set(keys))}")
digest = hashlib.sha256(repr(sorted((r.get("url"), r.get("price")) for r in good)).encode()).hexdigest()
if digest == last_good["digest"]:
alerts.append("every value identical to the last run: check for a stale source")
return alerts, digest
if __name__ == "__main__":
run = [
{"url": "https://shop.example/a", "title": "Kettle", "price": 39.0, "currency": "USD"},
{"url": "https://shop.example/b", "title": "Toaster", "price": 54.0, "currency": "USD", "sale_price": 59.0},
{"url": "https://shop.example/c", "title": "", "price": None, "currency": "USD"},
{"url": "https://shop.example/a", "title": "Kettle", "price": 39.0, "currency": "USD"},
]
for r in run:
print(r["url"], check_record(r) or "ok")
alerts, digest = check_run(run, {"rows": 4, "price_fill": 1.0, "digest": ""})
print(alerts)
https://shop.example/a ok
https://shop.example/b ['sale price above list price']
https://shop.example/c ['title missing or wrong type', 'price missing or wrong type']
https://shop.example/a ok
['rows fell to 2 from 4', 'price fill rate fell to 75%', 'duplicate keys: 1']
Save the digest from each run that passes and pass it back in as last_good["digest"] next time. The thresholds (80% of rows, a five-point fall in fill rate) are starting values. Tune them per source once you know how much each one moves. For larger pipelines, pydantic handles record schemas, Great Expectations handles run-level expectations, and Spidermon does both for Scrapy spiders, with alerts to Slack or email.
Buyers who ask us about data quality ask the same three things. Are the checks per row or per collection? Who sets the thresholds? What quality work is still left for us? The answers depend on how you buy.
With a scraping API, you own all five levels. The API's job is to return the page. A good one lets you tell a block from a real page, and String's Web Access API bills only when it returns content. The record, run, coverage and accuracy checks are still yours to write.
With a managed data service, the provider should run the response, record and run checks before each delivery, at both row and collection level, and tell you what they are. You should still set what "complete" means for each source, because only you know which fields your model needs. Keep your own accuracy sample too. String's Bespoke Web Datasets are monitored and validated before delivery, with a forward deployed engineer accountable for data quality. For what else to ask a managed provider, see managed scraping service vs scraping API.
Checks tell you whether the data matches the source page. They cannot tell you whether the source page is true. If a property site lists a unit as vacant after it is let, a perfect scraper stores a wrong vacancy. For figures that feed an investment or pricing decision, compare the series with an independent reference, such as the company's reported numbers, before you trust its level. And a success rate from any vendor, ours included, measures delivery, not correctness. Ask how the vendor defines a success before you compare two numbers.
Validate it at three levels. Check each response for content only the real page contains, each record against a schema with range and cross-field rules, and each run against the last run that passed. Then compare a random sample of records with the live pages by hand, because automated checks cannot see a value that is wrong but plausible.
Watch the output, not the process. Alert when the row count, a field's fill rate or the distinct key count moves away from the last good run for that source. Alert when every value is identical to the last run, which can mean a stale source. A scraper that runs without errors can still store block pages and empty fields.
Both. Row-level checks catch a missing price or a wrong currency in one record. Collection-level checks catch what no single row shows: a drop in row count, a fall in fill rate, duplicates, or a run where nothing changed. A pipeline that runs only one kind misses the failures of the other.
Give each record a stable key, such as the canonical URL or the site's own item ID, and upsert on that key instead of appending. Count duplicate keys on every run as a quality signal, because a jump in duplicates often means pagination broke. When the same item appears on several sites, match on a normalized identifier and keep the source of each copy.
A provider's success rate usually measures delivery: a 2xx status, sometimes with a check that the page loaded. It does not check that each field you extract is right. In our September 16, 2026 benchmark run, 7.7% of 2xx responses did not contain the expected page content. Ask the provider how it defines success, and run your own record and accuracy checks.
Match known challenge signatures first: vendor headers, challenge page titles, redirects to a login or consent path. A response with a block signature is a block, and a retry on a different route may fix it. A response with no block signature that fails your content marker is often a layout change, and the parser needs a fix.
The data buyer sets what complete means: which fields are required, how much coverage is enough, and how fresh the data must be. The collector sets the checks that enforce it and reports each failure. Review the first run of a new source by hand to set the expected shape, then let automated checks compare every later run with it.
Enough that every source and every page type appears in the sample on every run. Draw the records at random, compare each one with the live page, and record the error rate per source. A falling error rate shows a source is stable. A rising rate on one source is an early warning of a change on that site.
official_results/benchmark-2026-09-16T01-03-47-074Z.json in the public harness. Counts of attempts whose errorMessage is "Missing expected text" (2xx without the page marker), per provider and per target. Pass rule: validateResponse in src/check.ts.