NewLaunching String Web Access APIRead the manifesto →
← Blog

How to scrape Reddit in 2026

Bruce Magness · September 30, 2026
Part of Best Web Scraping APIs in 2026: 16 Tools Compared

Last updated: September 30, 2026. Benchmark figures come from the September 16, 2026 run of the Web Data Frontier Benchmark. Every script on this page was run on September 30, 2026, and the output shown is what it printed.

To scrape Reddit in 2026, add .json to a subreddit or thread URL and read the JSON Reddit returns. You do not need to parse the HTML. A listing such as https://www.reddit.com/r/webscraping/new/.json?limit=100 returns up to 100 posts with title, score, comment count and permalink, and an after cursor for the next page. A thread URL with .json returns the post and its whole comment tree. The hard part is getting Reddit to answer. On September 30, 2026, a plain Python requests call from a home connection got a JavaScript challenge on the subreddit page, a 403 "You've been blocked by network security" page on the .json URL, and a redirect to the login page on old.reddit.com. The Python below returned 200 posts in two requests through the String Web Access API, and the TypeScript returned a 47-comment thread in 1.5 seconds.

This page is part of our series on the best web scraping APIs, tested.

TL;DR

  • Reddit runs its own bot wall. The benchmark labels it "custom". A blocked .json request gets HTTP 403 and a page that says "You've been blocked by network security."
  • On September 30, 2026, plain requests from a home IP got no posts from any of the three URLs we tried. A Chrome User-Agent changed nothing.
  • In the September 16, 2026 benchmark, 10 of 16 scraping APIs returned reddit.com on all five attempts. Firecrawl, Context.dev and ZenRows returned it on none. Reddit ranked 65th of 100 sites by difficulty.
  • String returned Reddit on all five attempts, in 1.6 seconds on average.
  • Every Reddit request we logged through String billed as a premium fetch: $3.00 per 1,000 on Starter and $2.00 on Growth. One listing request carries up to 100 posts.

What people who scrape Reddit run into

The complaint in 2025 and 2026 is the same block page, in tool after tool. A gallery-dl user wrote in December 2025 that the Reddit extractor now stopped with "You've been blocked by network security", and added: "I can load the Reddit website perfectly in the browser, it is only when I use the program where it errors" (gallery-dl issue #8641, December 3, 2025). A Linkwarden user reported in February 2026 that capturing a Reddit page "only yields" that same message (Linkwarden issue #1611, February 17, 2026). The maintainer of content-core logged it as a feature request in April 2026: "Reddit blocks generic URL scraping with a 403 Forbidden response and an anti-bot challenge page" (content-core issue #35, April 5, 2026).

The first report has a detail worth keeping. The same user came back to say the tool worked again once they removed their cookies. A block like this scores the whole request, so a fix that works one week can stop working the next. Reddit's supported route is its Data API, which needs an OAuth app registered to a Reddit account and is governed by the Data API Terms. If your use fits those terms and the API's limits, use the API.

Why a plain request to Reddit fails

We sent three requests on September 30, 2026, from a home internet connection outside the US, each with Python's default User-Agent and again with a Chrome User-Agent. The first URL is the benchmark's Reddit URL.

What we tried, September 30, 2026 Default User-Agent Chrome User-Agent
https://www.reddit.com/r/webscraping/ HTTP 200, 8.4 KB, a js_challenge page, 0 posts Same
https://www.reddit.com/r/webscraping/.json HTTP 403, "You've been blocked by network security" Same
https://old.reddit.com/r/webscraping/ Redirected to /login/, 0 posts Same

The first row is the one to watch. It is an HTTP 200 with no data in it. A scraper that checks only the status code will record an empty subreddit rather than a block.

Through String, the same subreddit page came back as 545 KB of HTML with 13 shreddit-post elements, and the .json URL came back as a Listing with 25 posts, each in 2.4 seconds. The old.reddit.com URL failed through String too: an HTTP 502 from String after 63 seconds. The HTML route is not reliable either. Two days earlier, a scraping API (ours) returned a 362 KB r/fishing page with an empty feed, and the HTML route has no status code that tells you so. We use the .json route in both scripts below, because a Listing either has posts in it or is not a Listing.

Ways to get Reddit data, compared

Method Posts and comments What you maintain
Plain requests to .json No, 403 from a home IP in our test Nothing, and nothing works
Reddit's Data API (PRAW or direct OAuth) Yes, within the API's limits An OAuth app, a Reddit account, and the Data API Terms
Headless browser (Playwright, Selenium) Yes, slower A browser fleet, and HTML that Reddit changes
A prebuilt Reddit scraper from a marketplace such as Apify Yes An account there, and a scraper someone else keeps working
A fetch API such as the String Web Access API Yes, through .json Your parser, which is the JSON walk below

What 16 scraping APIs did with reddit.com

The Web Data Frontier Benchmark sends the same 100 URLs to 16 scraping APIs, five attempts each. The Reddit URL is the r/webscraping subreddit page. The harness and targets are open source, and so is the raw result file for the September 16 run.

API reddit.com, 5 attempts Average time on reddit.com Overall success, 100 sites
String 100% 1.6 s 97.0%
Scrapfly 100% 3.2 s 86.2%
ScraperAPI 100% 7.4 s 84.0%
Firecrawl 0% none passed 80.2%
Apify 100% 12.1 s 77.4%
Bright Data 100% 7.8 s 74.6%
ScrapingBee 100% 1.7 s 73.0%
Context.dev 0% none passed 72.0%
Oxylabs 100% 4.9 s 69.0%
Nimble 100% 3.6 s 68.6%
Zyte 100% 3.8 s 68.0%
Decodo 80% 28.2 s 50.6%
Scrapingdog 60% 32.8 s 45.6%
Browserbase 100% 1.5 s 41.4%
ZenRows 0% none passed 41.2%
ScrapingAnt 80% 4.3 s 36.4%

Ten of 16 APIs returned reddit.com on all five attempts: 61 passes out of 80 requests. The average API passed 76.3% of attempts on Reddit, against 66.6% across all 100 sites. Reddit ranked 65th of 100 by difficulty. Three APIs that do well overall, Firecrawl, Context.dev and ZenRows, passed none.

Read this row with one caveat. The benchmark counts a pass when the response is a 2xx and the body contains the text "webscraping", compared without regard to case. The 8.4 KB challenge page we got on a plain request also contains that text. So a pass here means the API got a 2xx page that mentions the subreddit, and not always a page with posts in it. Our String runs above returned posts, and that is the check to run on any tool before you commit to it: count the posts, not the status codes.

How to scrape a subreddit with Python

Install requests, set your key, and run it with a subreddit name and a page count. It reads Reddit's Listing JSON and follows the after cursor.

"""Scrape a subreddit's posts with the String Web Access API.

    pip install requests
    export STRING_API_KEY=...
    python reddit_subreddit.py webscraping 2
"""
import json
import os
import sys

import requests

API = "https://request.usestring.ai/v1/fetch"


def fetch_json(url: str) -> dict:
    response = requests.post(
        API,
        headers={"Authorization": f"Bearer {os.environ['STRING_API_KEY']}"},
        json={"url": url, "countryCode": "US"},
        timeout=90,
    )
    response.raise_for_status()
    page = response.json()
    if page["statusCode"] != 200:
        raise RuntimeError(f"reddit.com returned {page['statusCode']} for {url}")
    data = page["data"]
    # A JSON body comes back already parsed; an HTML body (a block page) comes back as a string.
    if isinstance(data, str):
        try:
            data = json.loads(data)
        except json.JSONDecodeError:
            raise RuntimeError("HTTP 200 but the body is HTML, not a Listing: treat this as blocked")
    if data.get("kind") != "Listing":
        raise RuntimeError(f"expected a Listing, got {data.get('kind')!r}")
    return data["data"]


def subreddit(name: str, pages: int = 2, sort: str = "new") -> list[dict]:
    rows, after = [], None
    for n in range(1, pages + 1):
        url = f"https://www.reddit.com/r/{name}/{sort}/.json?limit=100"
        if after:
            url += f"&after={after}"
        listing = fetch_json(url)
        posts = [c["data"] for c in listing["children"] if c["kind"] == "t3"]
        print(f"  page {n}: {len(posts)} posts, next cursor {listing['after']}")
        rows += [
            {
                "id": p["id"],
                "created_utc": int(p["created_utc"]),
                "author": p["author"],
                "title": p["title"],
                "score": p["score"],
                "comments": p["num_comments"],
                "url": "https://www.reddit.com" + p["permalink"],
            }
            for p in posts
        ]
        after = listing["after"]
        if not after:
            break
    return rows


if __name__ == "__main__":
    name = sys.argv[1] if len(sys.argv) > 1 else "webscraping"
    pages = int(sys.argv[2]) if len(sys.argv) > 2 else 2
    rows = subreddit(name, pages)
    print(f"r/{name}: {len(rows)} posts, {len({r['id'] for r in rows})} unique")
    for row in rows[:4]:
        print(f'{row["score"]:>4} pts {row["comments"]:>3} comments  {row["title"][:60]}')

What it printed on September 30, 2026, in 6.2 seconds:

  page 1: 100 posts, next cursor t3_1vmqxyi
  page 2: 100 posts, next cursor t3_1ucfe1r
r/webscraping: 200 posts, 200 unique
   3 pts   0 comments  My escalation ladder when a site blocks my basic requests
   2 pts   2 comments  Weekly Webscrapers - Hiring, FAQs, etc
   3 pts  17 comments  Would you consider selling your scraped data to AI agents?
   6 pts  10 comments  Is there a way to extract simple Yes/No answers from Google 

One detail in the String response matters here. When the page is JSON, String returns it already parsed in data. When it is HTML, data is a string. The script uses that difference as its block check: a string where a Listing should be is a challenge page. Each post in children carries more than the script keeps, including selftext, link_flair_text, upvote_ratio, over_18 and media URLs. Swap new for hot, top or rising to change the sort. We paged two pages; we did not test how deep the after cursor goes.

How to scrape a Reddit thread and its comments with TypeScript

No dependencies. Node 22.6 or later runs TypeScript directly (we ran it on Node 24), and so does Bun.

// Scrape one Reddit thread (the post and its comment tree) with the String Web Access API.
//   export STRING_API_KEY=...
//   node reddit_thread.ts <thread url>      (Node 22.6+; or: bun reddit_thread.ts <url>)

type Thing = { kind: string; data: Record<string, any> };
type Listing = { kind: "Listing"; data: { children: Thing[] } };

async function fetchJson(url: string): Promise<unknown> {
  const response = await fetch("https://request.usestring.ai/v1/fetch", {
    method: "POST",
    headers: { Authorization: `Bearer ${process.env.STRING_API_KEY}`, "Content-Type": "application/json" },
    body: JSON.stringify({ url, countryCode: "US" }),
  });
  if (!response.ok) throw new Error(`String returned ${response.status}: ${await response.text()}`);
  const page = (await response.json()) as { statusCode: number; data: unknown };
  if (page.statusCode !== 200) throw new Error(`reddit.com returned ${page.statusCode}`);
  // JSON bodies arrive parsed; a string here is an HTML page, which means a block, not a thread.
  if (typeof page.data === "string") throw new Error("HTTP 200 but the body is HTML: treat this as blocked");
  return page.data;
}

const thread = process.argv[2] ?? "https://www.reddit.com/r/webscraping/comments/1wowgnq/were_doing_web_scraping_wrong/";
const url = thread.replace(/\/?(\?.*)?$/, "/.json?limit=500&sort=top");
const [postListing, commentListing] = (await fetchJson(url)) as [Listing, Listing];

const post = postListing.data.children[0].data;
console.log(`${post.title}`);
console.log(`r/${post.subreddit} | ${post.score} points | ${post.num_comments} comments | ${new Date(post.created_utc * 1000).toISOString().slice(0, 10)}`);

// Walk the tree. "more" stubs are replies Reddit did not include in this response.
let comments = 0, maxDepth = 0, moreStubs = 0;
function walk(things: Thing[], depth: number) {
  for (const t of things) {
    if (t.kind === "more") { moreStubs += t.data.count ?? 0; continue; }
    if (t.kind !== "t1") continue;
    comments++;
    maxDepth = Math.max(maxDepth, depth);
    if (t.data.replies && typeof t.data.replies === "object") walk(t.data.replies.data.children, depth + 1);
  }
}
walk(commentListing.data.children, 0);
console.log(`${comments} comments parsed, deepest reply level ${maxDepth}, ${moreStubs} left in "more" stubs`);

const top = commentListing.data.children.find((t) => t.kind === "t1")?.data;
if (top) console.log(`top comment: ${top.score} points, ${String(top.body).length} characters, ${top.replies ? top.replies.data.children.length : 0} direct replies`);

We ran it on September 30, 2026, on Node 24 and on Bun. It returned in 1.5 seconds:

We're doing web scraping wrong
r/webscraping | 43 points | 52 comments | 2026-09-24
47 comments parsed, deepest reply level 9, 0 left in "more" stubs
top comment: 17 points, 619 characters, 3 direct replies

Reddit reported 52 comments and the tree held 47. The count includes comments that were removed or deleted, which do not appear in the tree. On a large thread, Reddit leaves part of the tree out and puts a more stub in its place with the IDs of the missing replies. The script counts those. This thread had none. To fill them in, fetch the permalink of the parent comment with .json; we did not test that on this run.

The failure mode: an HTTP 200 with no posts

Reddit's block comes in two shapes, and only one of them has an error code.

  • The 403 page. "You've been blocked by network security. If you think you've been blocked by mistake, file a ticket below." This is what a .json request gets when Reddit does not trust the client. It is easy to catch.
  • The challenge page, HTTP 200. The subreddit HTML URL returned an 8.4 KB page with a js_challenge script and no posts. A real subreddit page was 545 KB. A scraper that checks the status code logs this as a subreddit with nothing in it.

Three checks catch both:

  • Ask for JSON and check you got JSON. Both scripts raise an error when data is a string, because a Listing never is.
  • Check the kind. A subreddit or thread response is a Listing. Anything else is not the page you asked for.
  • Check the count. A listing with zero children from a subreddit that had posts yesterday is a block or a removed community. Log the URL and look at it.

The rest is Reddit's own shape. Comments that were removed or deleted still count toward num_comments, which is why our thread reported 52 and parsed 47. Do not treat that gap as a scraping error.

What it costs

Every Reddit request whose billing header we logged on September 30, 2026, the subreddit page and the .json listing, billed as a request-based fetch on a premium proxy. From String's pricing, that is $3.00 per 1,000 requests on Starter ($20 a month) and $2.00 per 1,000 on Growth ($100 a month). A request that fails is not billed.

Worked through from the runs above:

  • Subreddit monitoring. One listing request returns up to 100 posts. 100,000 posts a month is about 1,000 requests: $3 on Starter or $2 on Growth.
  • Threads with comments. One request per thread. 10,000 threads a month is $30 on Starter or $20 on Growth. A thread with more stubs costs one more request per stub you expand.

If your use fits Reddit's Data API terms, the API is the supported route and you should price it first.

FAQ

How to scrape Reddit posts?

Add .json to the subreddit URL, for example https://www.reddit.com/r/webscraping/new/.json?limit=100, and read data.children. Each child with kind "t3" is a post with title, score, num_comments, author and permalink. Pass the after value from one response as &after= on the next request to page.

How do I scrape Reddit comments?

Add .json to the thread URL. The response is two Listings: the first holds the post, the second holds the comment tree. Each comment is a "t1" with its replies nested under replies. The TypeScript script on this page walks the tree and returned 47 comments, nine levels deep, in 1.5 seconds on September 30, 2026.

Can I scrape Reddit without the API?

Yes, from public pages. In our test on September 30, 2026, a plain request from a home IP got a block or a challenge on every Reddit URL, and the same URLs returned posts through a fetch API. Reddit's supported route is the Data API, under its Data API Terms.

Why does Reddit say "You've been blocked by network security"?

It is Reddit's bot block. It returns that page with HTTP 403 when it does not trust the client, and developers report it from download tools, bookmark apps and scrapers through 2025 and 2026. A different User-Agent did not change it in our test. The request needs to come from a client and network Reddit accepts.

Does old.reddit.com still work for scraping?

Not in our test. On September 30, 2026, a logged-out request to https://old.reddit.com/r/webscraping/ was redirected to the login page, and the same URL failed through String with a 502 after 63 seconds. The .json route on www.reddit.com worked.

How many posts can I get from one Reddit request?

Up to 100, with limit=100 on a listing URL. Our Python script got 100 on each of two pages. The default without limit is 25, which is what the benchmark URL's .json returned.

Is there a free way to scrape Reddit?

Start with Reddit's Data API and read its terms for your use. PRAW is the usual Python client for it. For public pages without the API, nothing we tried from a home connection on September 30, 2026 returned posts.

Sources

Cheers,
String team

Get your API key →Explore the Web Access API
© 2026 StringEU and UK GDPR Article 27 representative — appointment verified by EuverifyBuilt in New York City 🗽 🍎