NewLaunching String Web Access APIRead the manifesto →
← Answers

How to get historical website data

String team · Updated September 30, 2026

Historical website data for a page nobody was collecting can come from only four places: a web archive such as the Wayback Machine or Common Crawl, history the site publishes itself, a vendor that was already collecting the page, or collection you start today. Archives are thin on deep pages such as product and listing pages. Across the 100 target URLs in String's Web Data Frontier Benchmark, the Wayback Machine holds a capture of the exact page in a median of 2 of the 12 months from September 2025 to August 2026, while 95 of the 100 home pages of those sites were captured every month. History rebuilt from an archive is also not point-in-time: it holds the days a crawler happened to visit, not a fixed series. For a backtest, buy history that someone collected on a schedule where it exists, and start your own collection now for everything else.

What the archives hold

On September 29 and 30, 2026 we queried the Wayback Machine's CDX index for the exact URL of each of the 100 targets in our benchmark (run of September 16, 2026), and counted the months from September 2025 to August 2026 with at least one capture. The home pages of the same sites had a capture in all 12 months for 95 of 100. The exact pages had captures in a median of 2 months: 22 had one every month, and 37 had none at all. Two of those 37 are on the blocked sites below. Of the other 35, 24 have no capture at any date. Some pages are younger than a year, such as a news live page from April 2026, which lowers their count.

Captures are not always the page. Of the 63 exact pages with any capture in the window, 21 had a month whose first capture was an HTTP error, and 9 had no HTTP 200 capture all year. For zillow.com's New York listings page, the archive holds captures in seven of the 12 months, and none of them returned HTTP 200. Two sites in the set, Yelp and Booking.com, are excluded from the Wayback Machine: the index answered our queries about them with "Blocked Site Error".

Older research found the same pattern. Ainsworth and colleagues sampled URLs from four sources and found that 35% to 90% had at least one archived copy, and no more than 31.3% were archived more than once a month (arXiv 1212.6177, the long version of a JCDL 2011 paper). Common Crawl published 12 crawls between September 2025 and August 2026, one a month, each gathered over about two weeks. Waiting also has a cost: Pew Research Center found that 38% of pages that existed in 2013 were gone by October 2023.

How to backfill a site

  1. Check the archive first. The CDX API lists every capture of a URL with its timestamp and HTTP status. Ask for exact matches, collapse by month, filter to status 200, then fetch the snapshots you need. We ran every index query for this page through String's Web Access API.
  2. Look for history the site keeps. Dated reviews, sold or expired listings, changelogs and archived posts carry their own dates. A sitemap's lastmod gives only the date of the last change, not earlier values (sitemaps.org).
  3. Ask a vendor for collected history, with each series' start date, its collection interval, and which periods were rebuilt from archives. String's Bespoke Web Datasets include 1000+ pre-built datasets with years of history. For what a managed provider should commit to, backfill included, see managed scraping service vs scraping API.
  4. Start collecting now. Save the raw page and a collection timestamp with every record, so the series is point-in-time from its first day.

Where archive history falls short

Archive gaps cannot be filled after the fact, and they are not random. In our test, most home pages had captures every month while most deep pages on the same sites did not, and on sites behind bot protection some captures were error pages. A 200 capture can still be a block page (how to detect one), so check each snapshot for the content you expect before you use it. Rebuilt history is also not point-in-time. The St. Louis Fed's ALFRED keeps every vintage of an economic series as it stood on each date; an archive keeps only what a crawler saw on the days it came. Label rebuilt periods in your data so a backtest does not treat them as live collection. For series already collected, String's Bespoke Web Datasets carry years of history, and String's Web Access API can start a scheduled, timestamped collection on the rest today.

FAQ

How do I get historical or archived web data at scale?

Query the Wayback Machine's CDX index to see which months exist for each URL, fetch those snapshots through a scraping API, and fill the rest from the sites' own dated records or a vendor's collected history. Expect gaps: in our test of 100 benchmark URLs, the exact page had a capture in a median of 2 of 12 months. Start scheduled collection for anything you will need later.

Can I backfill data from the Wayback Machine?

For home pages, often yes, because archives capture most of them every month. For deep pages such as a product or listing page, captures are sparse and some are error pages. Treat archive history as a partial record, and label it as rebuilt.

How far back does web-scraped data go?

As far back as someone collected it. A vendor's history starts when its collection started, which varies by series, and archives can extend some pages further back with gaps. Ask for the start date of each series before you plan a backtest. Buyers ask us this often: 26 of our 320 recorded sales calls with buyer questions raised it.

Is backfilled history point-in-time?

No. A point-in-time series records what a page showed on each date at a fixed interval. Archive history records what a crawler saw on the days it visited, so both the dates and the gaps depend on the crawler.

How often does the Wayback Machine capture a page?

It depends on the page. In our September 2026 test, 95 of 100 home pages had captures in every month of the prior year, while the exact pages we benchmark had captures in a median of 2 of 12 months.

Can history be rebuilt for any site?

Only where a copy exists. Records a site dates itself, such as reviews, commits or sold listings, can often be collected back in time. A value the site overwrites, such as a price or a stock level, exists in the past only if someone saved it.

Sources

  • String measurement: Wayback Machine CDX API, exact-URL queries for the 100 target URLs of the Web Data Frontier Benchmark (target list) and their home pages, window September 1, 2025 to August 31, 2026, queried September 29 and 30, 2026 through String's Web Access API.
  • Internet Archive, Wayback CDX Server API.
  • Common Crawl, crawl index list, checked September 29, 2026.
  • Scott G. Ainsworth, Ahmed AlSum, Hany SalahEldeen, Michele C. Weigle and Michael L. Nelson, "How Much of the Web Is Archived?", arXiv 1212.6177.
  • Pew Research Center, When Online Content Disappears, May 17, 2024.
  • Federal Reserve Bank of St. Louis, ALFRED.
  • Sitemaps.org, Sitemaps XML format.
  • String, Web Data Frontier Benchmark, run of September 16, 2026 (target list).
Get your API key →Explore the Web Access API
© 2026 StringEU and UK GDPR Article 27 representative — appointment verified by EuverifyBuilt in New York City 🗽 🍎