A scraping API returns clean structured JSON in one of three ways: a prebuilt endpoint for one site, AI extraction against a JSON Schema you send, or the JSON the page already carries. String does the last two in one API: its fetch returns the raw page with its embedded data intact, and a jsonSchema field returns only the fields you name. On October 5, 2026 we fetched the 100 sites in our benchmark: 64 of the 94 pages that loaded carried schema.org JSON-LD, and 29 carried a full Product record. On the 20 single-product pages where both methods ran, the page's own JSON-LD and AI extraction agreed on the price 16 times. Read the embedded JSON first and send a schema only for what is missing.
| Route | How it works | Fits | Watch for |
|---|---|---|---|
| Prebuilt endpoint for one site | The vendor maintains a parser and returns fixed fields | A few big sites you scrape every day | Only the sites and fields the vendor chose |
| AI extraction to your schema | You send a JSON Schema; a model reads the page and fills it | Any site, any fields | A per-page model fee, and a value the page does not hold can come back filled |
| The page's own embedded JSON | Parse JSON-LD, microdata or app state from the raw HTML | Product, listing, recipe, job and article pages | The site writes it for search engines, so it can lag the visible page |
ScraperAPI's structured data endpoints cover named sites such as Amazon, Google, Walmart, eBay and Redfin. String's site integrations return typed rows for Yelp and Google News. Schema extraction is offered by String, Firecrawl (a JSON format that adds 4 credits a page) and Scrapfly's Extraction API, among others. We have not benchmarked any vendor's extraction quality, so this page ranks none of them.
We fetched each of the 100 benchmark URLs once through String on October 5, 2026 and kept the raw body. 94 came back with a usable page.
__NEXT_DATA__ on 26.No JSON-LD block on any page failed to parse. Amazon, Target, Airbnb, Indeed and TikTok carried app state but no schema.org data, so their fields sit in site-specific JSON you map yourself.
A developer on Hacker News found the same thing on job boards: "I can cut down on DOM-based scraping a lot by just parsing schema.org-based JSON-LD" (chaosharmonic, September 13, 2026). Another warned the other way, that embedded data is "at worst it's a lie" when it drifts from what the page shows (blacksmith_tb). Both are right, and the test below shows where.
On the 27 pages with a Product and an Offer, we read the JSON-LD price and, seconds later, sent String a schema asking for name, price and currency. Five were search or listing pages (Zillow, Redfin, Realtor.com, Apartments.com, Zoopla), where a one-product schema is the wrong shape. That left 22 single-product pages.
| Result on 22 product pages | Pages |
|---|---|
| Both returned the same price | 16 |
| JSON-LD held the list price; the page showed a lower sale price, and extraction returned the sale price (Lululemon $148 vs $99, Ashley $1,399.99 vs $1,219.95) | 2 |
| The price sat only in the JSON-LD; extraction returned no price (Safeway) | 1 |
| The JSON-LD had no price; extraction found $3.39 (Kroger) | 1 |
| Page too large for extraction, so the request fell back to the raw page (AutoZone, AT&T) | 2 |
Across all 27 extraction requests the median time was 5.7 seconds, from 1.1 to 22.6. The two methods failed in different places, which is the case for running both. JSON-LD is free once you have the page, but it can carry the list price while the page sells at a discount. Extraction reads what a shopper sees, and on Safeway the price appeared only inside the page's JSON, not in its text.
Live pages also move. Kroger's page carried a JSON-LD price on one of four fetches we made that afternoon. Extraction returned a price on two of the other three.
On the listing pages, ask for an array. Redfin's one-product extraction returned the city's median sale price, $899,405, a real number from the page but not a listing. A schema with "listings": {"type": "array"} fits that page.
Fetch the page, read its JSON-LD, and send a schema only when the field you need is missing. This script does that. We ran it on October 5, 2026: Nike returned {"name": "KD19 \"Purple Stuff\" Basketball Shoes ...", "price": 124.97, "currency": "USD", "source": "json-ld"} with no extraction call.
"""Get a product as clean JSON: read the page's own JSON-LD first, ask for AI extraction only when it is missing.
pip install requests
export STRING_API_KEY=...
python product_json.py https://www.nike.com/t/kd19-purple-stuff-basketball-shoes-vrLfPVAT/IH1117-500
"""
import json
import os
import re
import sys
import requests
API = "https://request.usestring.ai/v1/fetch"
HEADERS = {"Authorization": f"Bearer {os.environ['STRING_API_KEY']}"}
LD_JSON = re.compile(r'<script[^>]+type=["\']?application/ld\+json["\']?[^>]*>(.*?)</script>', re.S | re.I)
SCHEMA = {
"type": "object",
"properties": {
"name": {"type": "string"},
"price": {"type": ["number", "null"]},
"currency": {"type": ["string", "null"], "pattern": "^[A-Z]{3}$"},
},
"required": ["name"],
}
def products(node):
"""Yield every schema.org Product object in a parsed JSON-LD tree."""
if isinstance(node, dict):
kind = node.get("@type")
if kind == "Product" or (isinstance(kind, list) and "Product" in kind):
yield node
for value in node.values():
yield from products(value)
elif isinstance(node, list):
for value in node:
yield from products(value)
def from_json_ld(html):
for block in LD_JSON.findall(html):
try:
tree = json.loads(block)
except json.JSONDecodeError:
continue
for product in products(tree):
offers = product.get("offers")
for offer in offers if isinstance(offers, list) else [offers]:
if isinstance(offer, dict) and offer.get("price") not in (None, ""):
return {"name": product.get("name"), "price": float(offer["price"]),
"currency": offer.get("priceCurrency"), "source": "json-ld"}
return None
def product_json(url):
page = requests.post(API, headers=HEADERS, json={"url": url}, timeout=120).json()
found = from_json_ld(page["data"]) if page["statusCode"] == 200 else None
if found:
return found
response = requests.post(API, headers=HEADERS, json={"url": url, "jsonSchema": SCHEMA}, timeout=180)
if response.headers.get("x-json-schema-applied") != "true":
return {"source": "none", "reason": response.headers.get("x-json-schema-fallback-reason")}
return {**response.json(), "source": "ai-extraction"}
if __name__ == "__main__":
print(json.dumps(product_json(sys.argv[1]), indent=2))
Three rules make the schema route cleaner. Mark only the fields you cannot do without as required: when extraction cannot fill one, String returns the raw page with extracted: false and a reason, not a half-filled object, per its structured extraction docs. Constrain free-text fields: in our run one currency field came back with an explanation written into it, which a pattern or enum stops. And check the x-json-schema-applied header before you trust the body. Our page on validating scraped data covers the checks to run after that.
String bills a fetch at the published rate, $0.20 per 1,000 standard requests on Growth and $0.30 on Starter, and nothing for a failed request. Extraction adds a fee scaled to page size and schema, charged whenever extraction runs, including a schema mismatch. Reading JSON-LD yourself adds no fee. String also returns Markdown for LLM input and raw HTML, from the same Web Access API.
Schema extraction reads the visible page, so a price that exists only in embedded JSON can come back empty, as it did on Safeway. Embedded JSON is written for search engines and can lag a sale. A prebuilt endpoint is the steadiest of the three on the sites it covers and absent everywhere else. For a list of sites that changes, embedded JSON first and a schema second returned a price on all 22 product pages in our test. Fetching the page at all is the first step, and on the Web Data Frontier Benchmark String returned content on 97.0% of requests across these 100 bot-protected sites, the September 16, 2026 run.
Several do, in three ways: prebuilt endpoints for named sites (ScraperAPI covers Amazon, Google, Walmart, eBay and Redfin), AI extraction to a JSON Schema you send (String, Firecrawl and Scrapfly offer it), and the JSON the page already carries. On 22 product pages from our benchmark, the page's own JSON-LD and String's extraction agreed on the price 16 times, so the strongest setup reads the embedded JSON first and sends a schema for what is missing.
A prebuilt endpoint runs a parser the vendor maintains for that site. AI extraction sends the page and your schema to a model, which fills the fields and coerces the types. Embedded-data parsing reads JSON-LD, microdata or app state the site already put in its HTML. Only the last costs nothing beyond the fetch.
Mostly. On 22 product pages it matched the price a model read from the page 16 times. On 2 it held the list price while the page showed a sale price, so compare it with the visible price when discounts matter.
It can fill a field the page does not hold, which is why practitioners validate with fixed rules. One developer put it as "getting schema-valid JSON is trivial. Getting correct values is not" (r/Rag). Keep required short, allow null, constrain strings with a pattern or enum, and check values against the page's JSON-LD when it exists.
For product pages, start with JSON-LD: 27 of the 29 Product records in our sample carried an Offer with a price. Add schema extraction for pages without it and for sale prices. A prebuilt endpoint suits a few marketplaces scraped at high volume. Our best e-commerce scraping APIs page ranks the APIs that fetch these sites.
It depends on the vendor. Firecrawl's JSON format adds 4 credits to a 1-credit scrape. String adds a fee scaled to page size and schema on top of the fetch rate, charged whenever extraction runs. Reading JSON-LD from a page you already fetched adds nothing.
Yes. Ask for an array of items in the schema, not one product. On Redfin's city page, a one-product schema returned the city's median sale price instead of a listing. Listing pages also often carry an ItemList in JSON-LD or the full result set in app state, such as Zillow's __NEXT_DATA__.
String skips extraction, returns the raw page in a fallback envelope with the reason body_too_large, and charges no extraction fee. Two of our 22 product pages, AutoZone and AT&T, hit this. The JSON-LD on both still carried the price.
official_results/benchmark-2026-09-16T01-03-47-074Z.json in the public harness.