NewLaunching String Web Access APIRead the manifesto →
← Blog

How to scrape Zillow in 2026

Bruce Magness · September 29, 2026
Part of Best Web Scraping APIs in 2026: 16 Tools Compared

Last updated: September 29, 2026. Benchmark figures come from the September 16, 2026 run of the Web Data Frontier Benchmark. Every script on this page was run on September 29, 2026, and the output shown is what it printed.

To scrape Zillow in 2026, fetch the search or property page and read the JSON Zillow puts in the page's __NEXT_DATA__ script tag. You do not need to parse the HTML. A search page carries 41 listings with zpid, price, beds, baths, square feet and address, and a property page carries the full home record. The part that decides whether you get data is the request itself. On September 29, 2026, a plain Python requests.get got Zillow's "Access to this page has been denied" page on every URL we tried. A browser User-Agent was enough for search pages. It was not enough for property pages, which answered 403 until the request also carried a browser's TLS fingerprint. The Python below returned 82 Austin listings in two requests and 4.3 seconds through the String Web Access API, and the TypeScript returned one full property record in 2.0 seconds.

This page is part of our series on the best web scraping APIs, tested.

TL;DR

  • Zillow runs PerimeterX, now HUMAN Security, per the benchmark's own label for the site. A blocked request gets HTTP 403 and a 5.8 KB page with a px-captcha element and no data.
  • A default requests call was blocked on search and property pages alike. With a Chrome User-Agent, search pages came back, 41 listings each, across 22 pages in a row. Property pages still returned 403.
  • curl_cffi impersonating Chrome's TLS fingerprint returned both search and property pages from a home connection. We did not test it from a cloud server, and that is where people report Zillow blocks most.
  • In the September 16, 2026 benchmark, 14 of 16 scraping APIs returned zillow.com on all five attempts. Zillow ranked 97th of 100 sites by difficulty, so it sits among the easiest targets in the run.
  • String returned Zillow on all five attempts but was slow there: 22.5 seconds on average, the second slowest of the 16.
  • Every Zillow request we logged through String billed as a standard fetch: $0.30 per 1,000 on Starter and $0.20 on Growth.

What people who scrape Zillow run into

The question in most threads is simple. A marketing lead on r/WebScrapingInsider wrote that her team had been "manually pulling listing data from Zillow for competitive reports" and asked how hard scraping it is now (r/WebScrapingInsider, 2026). An older Stack Overflow question from 2017 asks the same thing for all homes in Los Angeles, and names the parser trap: Zillow's HTML uses "dynamic ids (change every refresh)" (Stack Overflow, Chris Unice, October 7, 2017). Both problems go away when you read the embedded JSON rather than the markup.

The other pattern in the threads is code that works on a laptop and fails on a server. A developer on Google App Engine reported that Zillow calls worked "on my localhost" and on two hosted servers but failed in production on App Engine with a CAPTCHA (Stack Overflow, Kenny Wyland, January 20, 2017). That report is old. The shape still matches what bot walls do today: they score the network your request comes from, and cloud ranges score badly.

Why a plain request to Zillow fails

We sent three kinds of request to two search pages and three property pages on September 29, 2026, from a home internet connection. One of the search pages is the benchmark's Zillow URL, https://www.zillow.com/new-york-ny/.

What we tried, September 29, 2026 Search page Property page
requests.get, default User-Agent HTTP 403, 5.8 KB, "Access to this page has been denied", 0 listings HTTP 403, same block page
requests.get, Chrome User-Agent HTTP 200, about 630 KB, 41 listings HTTP 403, same block page
curl_cffi, impersonating Chrome HTTP 200, about 630 KB, 41 listings HTTP 200, about 750 KB, full property record

The User-Agent string is what gets you past the first check on search pages. Property pages check more: they let through only the request whose TLS handshake looked like Chrome's. We then paged through the New York search with a Chrome User-Agent: 22 pages, 41 listings each, 902 unique listings, no block. A burst of 60 search requests in 51 seconds also came back clean, all 60 with listings.

There are three reasons to use a service anyway, from our own runs and the threads above. The property pages need a browser-grade TLS fingerprint, which requests cannot send. Our tests ran from a home IP, and the reports of Zillow blocks come from cloud servers, where a scraper runs in production. And a block costs you a parser failure at 2 a.m., when the job that worked for months starts returning 5.8 KB pages. If you are pulling a few hundred listings from your laptop, try curl_cffi first.

Ways to get Zillow data, compared

Method Search listings Property details What you maintain
Plain requests Only with a browser User-Agent No, 403 A header string, and nothing when it breaks
curl_cffi with Chrome impersonation Yes, from a home IP Yes, from a home IP Impersonation profiles as Chrome updates; untested from cloud IPs
Headless browser (Playwright, Selenium) Yes, slower Yes, slower A browser fleet and a stealth setup
Zillow Group's own data and APIs Through a licensed program Through a licensed program An application to Zillow and its terms
A prebuilt scraper, such as String's Zillow Actors on Apify Yes, about 41 per location Yes, by ZPID An Apify account
A fetch API such as the String Web Access API Yes Yes Your parser, which is the __NEXT_DATA__ read below

What 16 scraping APIs did with zillow.com

The Web Data Frontier Benchmark sends the same 100 URLs to 16 scraping APIs, five attempts each. The Zillow URL is the New York for-sale search page. The harness and targets are open source, and so is the raw result file for the September 16 run.

API zillow.com, 5 attempts Average time on zillow.com Overall success, 100 sites
String 100% 22.5 s 97.0%
Scrapfly 100% 3.1 s 86.2%
ScraperAPI 100% 4.8 s 84.0%
Firecrawl 100% 13.8 s 80.2%
Apify 100% 4.8 s 77.4%
Bright Data 100% 2.2 s 74.6%
ScrapingBee 100% 2.4 s 73.0%
Context.dev 100% 5.3 s 72.0%
Oxylabs 100% 4.1 s 69.0%
Nimble 100% 3.1 s 68.6%
Zyte 100% 1.4 s 68.0%
Decodo 100% 39.7 s 50.6%
Browserbase 100% 1.8 s 41.4%
ZenRows 100% 10.6 s 41.2%
ScrapingAnt 80% 2.4 s 36.4%
Scrapingdog 0% none passed 45.6%

Fourteen of 16 APIs returned zillow.com on all five attempts: 74 passes out of 80 requests. The average API passed 92.5% of attempts on Zillow, against 66.6% across all 100 sites. Zillow ranked 97th of 100 by difficulty. ScrapingAnt's one miss was an HTTP 423 from its own API. Scrapingdog returned HTTP 500 on all five attempts.

String's row needs its own caveat. It passed five of five, but its average time on Zillow was 22.5 seconds, and only Decodo was slower. Three of String's five attempts took about 1 to 3 seconds and two took about 23. If time to first listing matters, Zyte, Browserbase, Bright Data and ScrapingBee were all under 2.5 seconds on this page.

The pass check here is stricter than it looks. The benchmark counts a pass when the response is a 2xx and the body contains "New York NY Real Estate - New York NY Homes For Sale", which is the search page's own title. Zillow's block page does not contain that text, so a blocked request cannot pass. The check does not count listings, though, so a page that loaded its title but not its results would still pass.

Zillow is easy in this run. Real estate as a whole is not. Across the seven real-estate sites in the benchmark, realtor.com came back on all five attempts from 5 of 16 APIs and apartments.com from 7. If Zillow is the first of several portals you need, test the others before you pick a tool.

How to scrape Zillow with Python

Install requests, set your key, and run it with a Zillow location slug and a page count. It reads the listings out of __NEXT_DATA__ and follows Zillow's own /2_p/ page URLs.

"""Scrape Zillow search results (for sale) with the String Web Access API.

    pip install requests
    export STRING_API_KEY=...
    python zillow_search.py "austin-tx" 2
"""
import json
import os
import re
import sys

import requests

API = "https://request.usestring.ai/v1/fetch"


def fetch(url: str) -> str:
    response = requests.post(
        API,
        headers={"Authorization": f"Bearer {os.environ['STRING_API_KEY']}"},
        json={"url": url, "countryCode": "US"},
        timeout=90,
    )
    response.raise_for_status()
    page = response.json()
    if page["statusCode"] != 200:
        raise RuntimeError(f"zillow.com returned {page['statusCode']} for {url}")
    return page["data"]


def listings(html: str) -> tuple[list[dict], int]:
    """Zillow ships the results as JSON in the page's __NEXT_DATA__ script tag."""
    m = re.search(r'<script id="__NEXT_DATA__"[^>]*>(.*?)</script>', html, re.S)
    if not m:
        # HTTP 200 without the data block: a challenge page, not an empty search.
        raise RuntimeError("HTTP 200 but no __NEXT_DATA__ in the page: treat this as blocked")
    state = json.loads(m.group(1))["props"]["pageProps"]["searchPageState"]
    results = state["cat1"]["searchResults"]["listResults"]
    return results, state["cat1"]["searchList"]["totalResultCount"]


def search(slug: str, pages: int = 2) -> list[dict]:
    rows, seen = [], set()
    for n in range(1, pages + 1):
        url = f"https://www.zillow.com/{slug}/" + (f"{n}_p/" if n > 1 else "")
        results, total = listings(fetch(url))
        new = [r for r in results if r["zpid"] not in seen]
        seen.update(r["zpid"] for r in new)
        print(f"  page {n}: {len(results)} listings, {len(new)} new, {total} in this search")
        for r in new:
            rows.append({
                "zpid": r["zpid"],
                "address": r.get("address"),
                "price": r.get("unformattedPrice"),
                "beds": r.get("beds"),
                "baths": r.get("baths"),
                "sqft": r.get("area"),
                "url": r.get("detailUrl"),
            })
    return rows


if __name__ == "__main__":
    slug = sys.argv[1] if len(sys.argv) > 1 else "austin-tx"
    pages = int(sys.argv[2]) if len(sys.argv) > 2 else 2
    rows = search(slug, pages)
    print(f"{slug}: {len(rows)} listings")
    for row in rows[:4]:
        print(f'{row["price"]:>9} {row["beds"]} bd {row["baths"]} ba {row["sqft"]} sqft  {row["address"]}')

What it printed on September 29, 2026, in 4.3 seconds:

  page 1: 41 listings, 41 new, 5700 in this search
  page 2: 41 listings, 41 new, 5700 in this search
austin-tx: 82 listings
   395000 3 bd 2 ba 1110 sqft  500 Denson Dr, Austin, TX 78752
   219000 3 bd 3 ba 1321 sqft  4619 Inicio Ln, Austin, TX 78725
   185000 1 bd 0 ba 306 sqft  404 Primrose St, Austin, TX 78753
   345000 3 bd 2 ba 1299 sqft  513 Shant St, Austin, TX 78748

Each entry in listResults carries more than the script keeps, including latLong, statusType, imgSrc and the broker name. Zillow reports 5,700 homes for this search but serves 41 per page, and the New York search reported 20 pages. To cover a whole city, split it by ZIP code or price band so each search stays inside the pages Zillow will serve.

How to scrape a Zillow property page with TypeScript

No dependencies. Node 22.6 or later runs TypeScript directly (we ran it on Node 24), and so does Bun.

// Scrape one Zillow property page (homedetails) with the String Web Access API.
//   export STRING_API_KEY=...
//   node zillow_detail.ts <homedetails url>      (Node 22.6+; or: bun zillow_detail.ts <url>)

type FetchEnvelope = { statusCode: number; data: string };

async function fetchPage(url: string): Promise<string> {
  const response = await fetch("https://request.usestring.ai/v1/fetch", {
    method: "POST",
    headers: { Authorization: `Bearer ${process.env.STRING_API_KEY}`, "Content-Type": "application/json" },
    body: JSON.stringify({ url, countryCode: "US" }),
  });
  if (!response.ok) throw new Error(`String returned ${response.status}: ${await response.text()}`);
  const page = (await response.json()) as FetchEnvelope;
  if (page.statusCode !== 200) throw new Error(`zillow.com returned ${page.statusCode}`);
  return page.data;
}

const url = process.argv[2] ?? "https://www.zillow.com/homedetails/500-Denson-Dr-Austin-TX-78752/29412534_zpid/";
const html = await fetchPage(url);

// The property record sits in __NEXT_DATA__ > props.pageProps.componentProps.gdpClientCache,
// which is itself a JSON string keyed by a GraphQL query name.
const nextData = html.match(/<script id="__NEXT_DATA__"[^>]*>([\s\S]*?)<\/script>/)?.[1];
if (!nextData) throw new Error("HTTP 200 but no __NEXT_DATA__ in the page: treat this as blocked");
const cache = JSON.parse(JSON.parse(nextData).props.pageProps.componentProps.gdpClientCache);
const home = Object.values(cache as Record<string, { property?: Record<string, any> }>).find((v) => v.property)?.property;
if (!home) throw new Error("page loaded but carried no property record");

console.log(`${home.address.streetAddress}, ${home.address.city}, ${home.address.state} ${home.address.zipcode}`);
console.log(`status ${home.homeStatus} | price ${home.price} | zestimate ${home.zestimate ?? "none"}`);
console.log(`${home.bedrooms} bd | ${home.bathrooms} ba | ${home.livingArea} sqft | built ${home.yearBuilt} | lot ${home.lotSize ?? home.lotAreaValue ?? "n/a"}`);
console.log(`${(home.description ?? "").slice(0, 90)}`);

We ran it on September 29, 2026, on Node 24 and again on Bun, with the same output. It returned in 2.0 seconds:

500 Denson Dr, Austin, TX 78752
status FOR_SALE | price 395000 | zestimate none
3 bd | 2 ba | 1110 sqft | built 1951 | lot 6699
Enjoy easy North Central Austin living, where morning walks lead straight to Kyoko or Benn

The property object holds far more than four lines: price history, tax history, schools, the listing agent, resoFacts with the MLS fields, and photos. This listing had no Zestimate, so the script prints "none" rather than failing.

The failure mode: a 403 that looks like a page

Zillow's block is honest about its status code, which makes it easier to catch than most. Every blocked request we sent came back HTTP 403 with a page of about 5.8 KB titled "Access to this page has been denied", carrying a px-captcha element. A real search page is about 600 KB and a property page about 600 to 750 KB.

Three checks catch it:

  • Check the status. A 403 from zillow.com is the block page. Retrying the same request from the same network gets the same page.
  • Check for __NEXT_DATA__. Both scripts raise an error when the tag is missing, so a challenge page never reaches your parser.
  • Check the count. A search page that parses but carries zero listResults is either a search with no homes or a partial page. Log it with the URL and look at a few by hand.

The quieter risk is the parser. Zillow changes the shape of __NEXT_DATA__ from time to time. When searchPageState or gdpClientCache moves, the scripts raise a KeyError or TypeError rather than write empty rows. That is the behavior you want.

What it costs

Both Zillow requests whose billing header we logged on September 29, 2026, one search page and one property page, billed as a standard fetch. From String's pricing, that is $0.30 per 1,000 requests on Starter ($20 a month) and $0.20 per 1,000 on Growth ($100 a month). A request that fails is not billed.

Worked through from the runs above:

  • Listings. One search page returns 41 listings. 100,000 listings a month is about 2,440 requests: under $1 in usage on either plan.
  • Property pages. One request per home. 100,000 property pages a month is $30 on Starter or $20 on Growth.

For a small, occasional pull, curl_cffi from your own machine costs nothing and worked in our test. The service earns its cost when the job runs on a schedule from a server.

FAQ

How to scrape Zillow listings?

Fetch the search page for a city or ZIP, for example https://www.zillow.com/austin-tx/, and parse the JSON in its __NEXT_DATA__ script tag. The listings are in searchPageState.cat1.searchResults.listResults, 41 per page, each with zpid, price, beds, baths, area and address. Add /2_p/, /3_p/ and so on for more pages.

How do I scrape Zillow with Python?

Send the Zillow URL to a fetch API with requests, then load the __NEXT_DATA__ JSON with the standard library's json module. The Python script on this page does that and returned 82 Austin listings in two requests and 4.3 seconds on September 29, 2026.

Why does Zillow return 403 "Access to this page has been denied"?

That is the PerimeterX (HUMAN Security) block page. In our test on September 29, 2026, a default requests User-Agent got it on every Zillow URL, and a Chrome User-Agent still got it on property pages. Property pages loaded only when the request also carried a browser TLS fingerprint.

Does Zillow have an API?

Zillow Group offers data and APIs through its developer programs, which are licensed and approved per use case. They are not an open API for pulling any listing. For public listing pages, people scrape the site or use a scraper built on it.

How many listings can I scrape from one Zillow search?

Zillow serves 41 listings per page and caps the number of pages. The New York search reported 20 pages; Austin reported 5,700 homes in total. To go past the cap, split the area into smaller searches, such as by ZIP code or price range.

Can I scrape Zillow property details like the Zestimate and price history?

Yes. The property page carries them in __NEXT_DATA__ under componentProps.gdpClientCache, as a JSON string that holds a property object. The TypeScript script on this page reads address, status, price, Zestimate, beds, baths, living area, year built and lot size from it. Not every listing has a Zestimate.

Is there a free way to scrape Zillow?

For small jobs from your own computer, curl_cffi with Chrome impersonation returned both search and property pages in our test on September 29, 2026. We ran 22 search pages in a row without a block. We did not test it from a cloud server, where most reports of Zillow blocks come from.

Sources

Cheers,
String team

Get your API key →Explore the Web Access API
© 2026 StringEU and UK GDPR Article 27 representative — appointment verified by EuverifyBuilt in New York City 🗽 🍎