Historical website data for a page nobody was collecting can come from only four places: a web archive such as the Wayback Machine or Common Crawl, history the site publishes itself, a vendor that was already collecting the page, or collection you start today. Archives are thin on deep pages such as product and listing pages. Across the 100 target URLs in String's Web Data Frontier Benchmark, the Wayback Machine holds a capture of the exact page in a median of 2 of the 12 months from September 2025 to August 2026, while 95 of the 100 home pages of those sites were captured every month. History rebuilt from an archive is also not point-in-time: it holds the days a crawler happened to visit, not a fixed series. For a backtest, buy history that someone collected on a schedule where it exists, and start your own collection now for everything else.
On September 29 and 30, 2026 we queried the Wayback Machine's CDX index for the exact URL of each of the 100 targets in our benchmark (run of September 16, 2026), and counted the months from September 2025 to August 2026 with at least one capture. The home pages of the same sites had a capture in all 12 months for 95 of 100. The exact pages had captures in a median of 2 months: 22 had one every month, and 37 had none at all. Two of those 37 are on the blocked sites below. Of the other 35, 24 have no capture at any date. Some pages are younger than a year, such as a news live page from April 2026, which lowers their count.
Captures are not always the page. Of the 63 exact pages with any capture in the window, 21 had a month whose first capture was an HTTP error, and 9 had no HTTP 200 capture all year. For zillow.com's New York listings page, the archive holds captures in seven of the 12 months, and none of them returned HTTP 200. Two sites in the set, Yelp and Booking.com, are excluded from the Wayback Machine: the index answered our queries about them with "Blocked Site Error".
Older research found the same pattern. Ainsworth and colleagues sampled URLs from four sources and found that 35% to 90% had at least one archived copy, and no more than 31.3% were archived more than once a month (arXiv 1212.6177, the long version of a JCDL 2011 paper). Common Crawl published 12 crawls between September 2025 and August 2026, one a month, each gathered over about two weeks. Waiting also has a cost: Pew Research Center found that 38% of pages that existed in 2013 were gone by October 2023.
lastmod gives only the date of the last change, not earlier values (sitemaps.org).Archive gaps cannot be filled after the fact, and they are not random. In our test, most home pages had captures every month while most deep pages on the same sites did not, and on sites behind bot protection some captures were error pages. A 200 capture can still be a block page (how to detect one), so check each snapshot for the content you expect before you use it. Rebuilt history is also not point-in-time. The St. Louis Fed's ALFRED keeps every vintage of an economic series as it stood on each date; an archive keeps only what a crawler saw on the days it came. Label rebuilt periods in your data so a backtest does not treat them as live collection. For series already collected, String's Bespoke Web Datasets carry years of history, and String's Web Access API can start a scheduled, timestamped collection on the rest today.
Query the Wayback Machine's CDX index to see which months exist for each URL, fetch those snapshots through a scraping API, and fill the rest from the sites' own dated records or a vendor's collected history. Expect gaps: in our test of 100 benchmark URLs, the exact page had a capture in a median of 2 of 12 months. Start scheduled collection for anything you will need later.
For home pages, often yes, because archives capture most of them every month. For deep pages such as a product or listing page, captures are sparse and some are error pages. Treat archive history as a partial record, and label it as rebuilt.
As far back as someone collected it. A vendor's history starts when its collection started, which varies by series, and archives can extend some pages further back with gaps. Ask for the start date of each series before you plan a backtest. Buyers ask us this often: 26 of our 320 recorded sales calls with buyer questions raised it.
No. A point-in-time series records what a page showed on each date at a fixed interval. Archive history records what a crawler saw on the days it visited, so both the dates and the gaps depend on the crawler.
It depends on the page. In our September 2026 test, 95 of 100 home pages had captures in every month of the prior year, while the exact pages we benchmark had captures in a median of 2 of 12 months.
Only where a copy exists. Records a site dates itself, such as reviews, commits or sold listings, can often be collected back in time. A value the site overwrites, such as a price or a stock level, exists in the past only if someone saved it.