NewLaunching String Web Access APIRead the manifesto →
← Blog

How to scrape flight prices with Python

Bruce Magness · October 7, 2026
Part of Best Web Scraping APIs in 2026: 16 Tools Compared

Last updated: October 7, 2026. Benchmark figures come from the September 16, 2026 run of the Web Data Frontier Benchmark. The script on this page ran on October 7, 2026, and the output shown is what it printed.

To scrape flight prices with Python, fetch a Google Flights search URL and read the fare cards in the HTML. Each card carries a plain-English label with the price, airline, stops, times and duration, so one regular expression parses it and no browser is needed. On October 7, 2026, the script below returned 50 flights for San Francisco to New York on two dates in 5.2 seconds, cheapest $209. Airline and travel sites are the harder part. On four flight sites in our benchmark, 12 of 16 scraping APIs failed every attempt on at least one of them.

This page is part of our series on the best web scraping APIs, tested.

On this page: Where String fits · Step 1: build the search URL · Step 2: fetch the page · Step 3: parse the fare cards · Step 4: track dates and save a CSV · How to avoid getting blocked on flight sites · How String gets flight pages · FAQ

Where String fits

String is a fit for one job here: getting the page back when a travel site blocks your scraper.

  • For: developers and small teams that track fares on a fixed list of routes and dates, and want one API call per search instead of proxies, browsers and retries.
  • Not for: booking flows, pages behind a login you do not own, or a business that needs a contracted fare feed from a GDS.
  • Price: you pay only when content comes back. On the Starter plan ($20 a month), a plain fetch is $0.30 per 1,000 on standard proxies and $3.00 on premium; a browser fetch is $1.50 and $6.00. The first 5,000 standard requests are free (pricing).
  • Integrations: a REST API you call from Python or any language, and an MCP server for agents.
  • Limits: 60 requests a second per account with a 3,600-request burst, and no cap on concurrency.

Where flight prices come from

Fares come from three places: airline sites, online travel agencies, and metasearch sites.

Airlines sell their own seats on sites such as aa.com and delta.com. Online travel agencies such as Expedia resell them. Metasearch sites such as Google Flights, Kayak and Skyscanner show fares from many airlines and agencies on one page, which makes them the cheapest place to collect a route.

The free official routes have mostly closed. Google shut down its QPX Express flight API on April 10, 2018 (TechCrunch). Amadeus decommissioned its self-service developer portal on July 17, 2026, and its site now serves the Enterprise API portal only (Amadeus for Developers).

That gap is why people scrape. One developer on r/webscraping put it this way in March 2026:

"The first issue is finding an api for flight fares. Last time I was looking for such service, I found nothing free or accessible for solo/hobby developers." Puzzleheaded-War3790, r/webscraping, March 29, 2026

Scraping is not always the right answer. In the same thread, another commenter told the poster, who needed one trip planned, that searching by hand would be faster:

"i seriously doubt scraping is the solution here. to build a scraping solution you will need to spend probably 10-20 hours and i bet you can manually do it way faster for this one-off need" Objectdotuser, r/webscraping, March 14, 2026

That holds for one trip. A script pays off when you check the same routes every day.

Step 1: build the search URL

Google Flights accepts a plain-language query in the q parameter.

The script builds that query from an origin, a destination and a date, and pins the language, country and currency so every run is comparable:

from urllib.parse import quote

def search_url(origin: str, dest: str, date: str) -> str:
    q = f"Flights to {dest} from {origin} on {date} oneway"
    return f"https://www.google.com/travel/flights?q={quote(q)}&hl=en&gl=us&curr=USD"

Fix curr and gl on every request. Fares change with the country and currency of the search, so a run that drifts between them compares different prices.

Step 2: fetch the page

The fetch is one POST to String's fetch endpoint with the target URL.

import os
import requests

API = "https://request.usestring.ai/v1/fetch"

def fetch(url: str) -> str:
    response = requests.post(
        API,
        headers={"Authorization": f"Bearer {os.environ['STRING_API_KEY']}"},
        json={"url": url, "countryCode": "US"},
        timeout=90,
    )
    response.raise_for_status()
    page = response.json()
    if page["statusCode"] != 200:
        raise RuntimeError(f"google.com returned {page['statusCode']} for {url}")
    return page["data"]

statusCode is the site's own status, separate from the API's. countryCode: "US" routes the request through a US exit, which keeps the fares in line with gl=us.

Google Flights also answered a plain requests.get from a US home connection on October 7, 2026, with all 48 fare-card labels. At one search a day you may not need an API at all. Volume is where Google pushes back. A developer planning 6,500 calendar requests per run was told in r/webscraping that Google would block them, and that they would need to pay for API access or proxies (thread).

Run this with a free key: the first 5,000 standard requests cost nothing.

Step 3: parse the fare cards

Every fare card has an aria-label that states the whole result in one sentence.

This is one label from the October 7 run, as served:

From 209 US dollars. Nonstop flight with Delta. Leaves San Francisco International Airport at 9:00 AM on Thursday, November 12 and arrives at John F. Kennedy International Airport at 5:33 PM on Thursday, November 12. Total duration 5 hr 33 min.

Accessibility labels change less often than class names, which Google generates and rotates. One regular expression reads it:

import re
from datetime import datetime, timezone

CARD = re.compile(
    r'aria-label="From ([\d,]+) US dollars\. (Nonstop|\d+ stops?) flight with ([^.]+)\. '
    r"Leaves (.+?) at ([\d:]+\s?[AP]M) on [^.]*? and arrives at (.+?) at ([\d:]+\s?[AP]M) on [^.]*?\. "
    r"Total duration ([^.]+)\."
)

def fares(origin: str, dest: str, date: str) -> list[dict]:
    html = fetch(search_url(origin, dest, date))
    rows, seen = [], set()
    for m in CARD.finditer(html):
        price, stops, airline, dep_ap, dep, arr_ap, arr, duration = m.groups()
        key = (airline, dep, arr, price)
        if key in seen:  # the page repeats some cards (best flights, then all flights)
            continue
        seen.add(key)
        rows.append({
            "scraped_at": datetime.now(timezone.utc).isoformat(timespec="seconds"),
            "origin": origin, "dest": dest, "date": date,
            "price_usd": int(price.replace(",", "")),
            "stops": 0 if stops == "Nonstop" else int(stops.split()[0]),
            "airline": airline, "departs": dep.replace(" ", " "), "arrives": arr.replace(" ", " "), "duration": duration,
        })
    if not rows:
        # A 200 with no fare cards is a block page, a consent page, or a layout change.
        raise RuntimeError(f"HTTP 200 but no fare cards for {origin}-{dest} on {date}")
    return sorted(rows, key=lambda r: r["price_usd"])

Two details matter. Google puts a narrow no-break space ( ) between the time and AM or PM, so the script swaps it for a normal space. And the page lists some flights twice, once under "best" and once in the full list, so the script drops repeats.

The empty-result check is the most useful line in the file. A block page, a consent screen and a redesign all arrive as HTTP 200, and without the check they become empty rows in your data. Our page on detecting a block page that returns HTTP 200 covers that failure in depth.

Step 4: track dates and save a CSV

Prices only mean something as a series, so the script takes several dates and appends every run to one CSV.

import csv
import sys

if __name__ == "__main__":
    args = sys.argv[1:]
    out = None
    if "--csv" in args:
        i = args.index("--csv")
        out = args[i + 1]
        args = args[:i] + args[i + 2:]
    origin, dest, *dates = args or ["SFO", "JFK", "2026-11-12"]
    all_rows = []
    for date in dates:
        rows = fares(origin, dest, date)
        all_rows += rows
        print(f"{origin}-{dest} {date}: {len(rows)} flights, cheapest ${rows[0]['price_usd']}")
        for r in rows[:5]:
            print(f"  ${r['price_usd']:>5}  {r['airline'][:28]:<28} {r['stops']} stop(s)  {r['departs']} -> {r['arrives']}  {r['duration']}")
    if out:
        new = not os.path.exists(out)
        with open(out, "a", newline="") as f:
            w = csv.DictWriter(f, fieldnames=list(all_rows[0]))
            if new:
                w.writeheader()
            w.writerows(all_rows)
        print(f"appended {len(all_rows)} rows to {out}")

Run it with pip install requests, export STRING_API_KEY=..., then:

python flight_prices.py SFO JFK 2026-11-12 2026-11-19 --csv fares.csv

This is what it printed on October 7, 2026, in 5.2 seconds:

SFO-JFK 2026-11-12: 24 flights, cheapest $209
  $  209  Delta                        0 stop(s)  9:00 AM -> 5:33 PM  5 hr 33 min
  $  209  Alaska                       0 stop(s)  1:50 PM -> 10:29 PM  5 hr 39 min
  $  209  JetBlue                      0 stop(s)  4:25 PM -> 12:56 AM  5 hr 31 min
  $  209  Delta                        0 stop(s)  7:00 AM -> 3:31 PM  5 hr 31 min
  $  209  Delta                        0 stop(s)  11:50 AM -> 8:20 PM  5 hr 30 min
SFO-JFK 2026-11-19: 26 flights, cheapest $242
  $  242  Delta                        1 stop(s)  10:38 AM -> 12:09 AM  10 hr 31 min
  $  254  American                     1 stop(s)  8:29 AM -> 7:00 PM  7 hr 31 min
  $  254  American                     1 stop(s)  11:57 AM -> 10:14 PM  7 hr 17 min
  $  269  Alaska                       1 stop(s)  10:39 AM -> 10:55 PM  9 hr 16 min
  $  319  Alaska                       0 stop(s)  7:00 AM -> 3:35 PM  5 hr 35 min
appended 50 rows to fares.csv

A long-haul search works the same way. Chicago to London on December 3, 2026 returned 15 flights, cheapest $608 on Air France and Delta. Put the command in cron once a day and the CSV becomes a fare history.

How to avoid getting blocked on flight sites

Flight sites are harder to scrape than the average site in our benchmark.

The September 16, 2026 run loaded four flight pages five times each through 16 scraping APIs: the aa.com home page, delta.com flight status, a Kayak route page and a Skyscanner route page. The field returned 190 of 320 loads (59.4%). Across all 100 pages in the run, the average page succeeded 66.6% of the time.

API aa.com delta.com kayak.com skyscanner.net Loads out of 20
String 5 5 5 5 20
Bright Data 4 5 5 5 19
ScraperAPI 3 5 5 5 18
Scrapfly 2 5 5 4 16
Firecrawl 0 5 5 5 15
Apify 5 5 5 0 15
Nimble 5 5 5 0 15
Context.dev 4 5 5 0 14
Scrapingdog 4 5 5 0 14
ScrapingAnt 0 5 5 4 14
Zyte 0 4 5 3 12
Oxylabs 5 5 0 0 10
Decodo 5 0 0 0 5
ScrapingBee 3 0 0 0 3
Browserbase 0 0 0 0 0
ZenRows 0 0 0 0 0

Successful loads out of five per site, read from the per-target cells of the September 16, 2026 run. A load passes only when the API returns a 2xx status and the body contains text from the real page, so a challenge page served at 200 counts as a failure. The harness and every adapter are open source.

Skyscanner was the hardest of the four. Nine of 16 APIs got nothing from it on any attempt. Only String, Bright Data, Firecrawl and ScraperAPI returned it five times out of five.

Each site failed in its own way when we fetched it with plain Python on October 7, 2026:

Site Plain requests from a home connection Through String
aa.com 403 "Access Denied", 379 bytes 200, 86 KB, plain fetch, 1.3 s
delta.com 200, 4 KB shell, content needs JavaScript 200, 303 KB, browser fetch, 16.6 s
kayak.com route page 200, full page 200, full page, plain fetch, 1.8 s
skyscanner.net route page 200, 708-byte shell 200, 645 KB, browser fetch, 34.4 s

Our benchmark groups aa.com and delta.com under Akamai Bot Manager and Skyscanner under PerimeterX. Our page on getting past Akamai Bot Manager explains the Akamai checks. Kayak's route page came back to a plain client, but its fare results did not: the HTML for a dated Kayak search held no fares until a browser ran the page's JavaScript (we waited 12 seconds), and then the first price on the page was a paid placement.

The rules that keep a flight scraper running:

  • Prefer metasearch over airline sites. One Google Flights page carries every carrier on the route, so you make fewer requests and meet fewer walls.
  • Test the body, not the status. Count fare cards on every response, as Step 3 does, and treat zero as a failure.
  • Hold the country and currency fixed. Fares vary by market. One commenter on r/webscraping warned that "pricing may display differently when using a proxy (different region, different demographic/profile, etc)" (--Adam, March 14, 2026). Use one exit country per series.
  • Pace your requests. Our page on handling rate limits covers backoff and how a 429 can hide inside a 200.
  • Use a browser only where the fares need one. The browser fetches in the table above took 13 to 26 times as long as the plain fetch of aa.com.

How String gets flight pages

String is the most reliable web scraping API, bypassing anti-bot, and you only pay when content comes back. For flight pages, that means:

  • One call per search URL. You send POST /v1/fetch with the URL, and String picks a plain request, a premium residential proxy or a real browser (fetch docs).
  • The billed path is in the response. The x-billed-request-type header named request_standard for aa.com and Google Flights and browser_premium for delta.com and Skyscanner on October 7, 2026.
  • The site's own status comes back. statusCode in the response body is what the airline or travel site answered, so a real 404 for a retired route stays a 404.
  • String returned all four flight pages on every attempt. In the September 16, 2026 run, String was the only one of 16 APIs to return aa.com, delta.com, kayak.com and skyscanner.net five times out of five.
  • A failed fetch is not billed. String charges only when content comes back.
  • Fare tracking is cheap at plain-fetch rates. Twenty routes on 30 dates a day is 18,000 requests a month, which is $5.40 at the Starter plan's standard fetch rate of $0.30 per 1,000.

Is it legal to scrape flight prices?

Reading public fare pages is legal in many places, but airlines have sued over it.

Ryanair sued Booking.com under the US Computer Fraud and Abuse Act over screen scraping of ryanair.com. A Delaware jury found Booking.com liable in July 2024 (Reuters). The court overturned the verdict in January 2025 because Ryanair had not proved $5,000 in losses, and the case went to the Third Circuit on appeal (Reporters Committee for Freedom of the Press). The same ruling found that the CFAA can bar access to pages that need a free account once the site owner has sent a cease-and-desist letter. Keep to public search pages, read each site's terms, and see whether web scraping is legal. This is not legal advice.

Start with the routes you check by hand today. Run the script above on them with a free key and compare the CSV a week from now.

FAQ

How do I scrape flight prices with Python?

Fetch a Google Flights search URL with requests and parse the aria-label on each fare card, which states price, airline, stops, times and duration in one sentence. One regular expression reads it. Route the fetch through a scraping API when Google starts blocking your IP, and check every response for at least one fare card.

Does Google Flights have an API?

No public one. Google shut down its QPX Express flight API on April 10, 2018. Developers now read the Google Flights pages directly or buy fare data from a vendor.

Is there a free API for flight prices?

Few remain for individual developers. Amadeus decommissioned its self-service portal, which many hobby projects used, on July 17, 2026. Scraping a metasearch page is now the common free option at low volume; at higher volume you pay either for proxies or for a scraping API.

Why does my flight price scraper get blocked?

Airline and travel sites run bot managers. In our September 16, 2026 benchmark, aa.com and delta.com sat behind Akamai and Skyscanner behind PerimeterX. A plain Python client got a 403 from aa.com and an empty shell from Skyscanner on October 7, 2026. Fix the IP, the TLS fingerprint and JavaScript rendering, or use an API that does.

Which scraping API works best on flight sites?

In String's September 16, 2026 run on four flight pages, String returned 20 of 20 loads, Bright Data 19 and ScraperAPI 18. Nine of 16 APIs got nothing from Skyscanner on any attempt. We build String, so check the per-target data on the benchmark page yourself.

Can I scrape Kayak or Skyscanner?

Both pages came back through String on October 7, 2026, but their fares need a browser. Kayak's dated search HTML held no fares until its JavaScript ran, and Skyscanner served a 708-byte shell to a plain client. Google Flights put its fares in the first HTML, so it is the cheaper source.

Do flight prices change when you scrape through a proxy?

They can. Fares and currency vary by the country of the search, so a proxy in another country can show a different price. Fix the exit country and the gl and curr parameters for every request in a series.

How often should I scrape flight prices?

Once a day per route and date is enough to build a fare history, and it keeps the request count low. Each extra check per day adds cost and raises the chance of a block. Prices are dynamic, so a commenter on r/webscraping noted that you "would need to scrape frequently" to catch changes; set the rate by how fast you must react.

Sources

Cheers,
String team

Get your API key →Explore the Web Access API
© 2026 StringEU and UK GDPR Article 27 representative — appointment verified by EuverifyBuilt in New York City 🗽 🍎