NewLaunching String Web Access APIRead the manifesto →
← Blog

How to scrape X (Twitter) in 2026

Bruce Magness · September 30, 2026
Part of Best Web Scraping APIs in 2026: 16 Tools Compared

Last updated: September 30, 2026. Benchmark figures come from the September 16, 2026 run of the Web Data Frontier Benchmark. Every script on this page was run on September 30, 2026, and the output shown is what it printed.

To scrape X (Twitter) in 2026, fetch the public profile or post page and read the posts out of the HTML. X now puts them in the page it serves to logged-out visitors, so you do not need to run JavaScript. What you can get is what X shows a logged-out visitor: a profile's bio and follower count, its five latest posts with likes and views, and a single post with its first few replies. Search, full timelines and full reply threads sit behind the login wall. On September 30, 2026, the Python below returned NASA's profile and five posts with likes and views in 2.1 seconds through the String Web Access API, and the TypeScript returned one post with its reply, repost, like and view counts in 1.7 seconds. A search URL returned a 17 KB empty app shell.

This page is part of our series on the best web scraping APIs, tested. It covers X in full, and Instagram and TikTok in the section near the end.

TL;DR

  • X runs its own bot defense, labeled "custom" in the benchmark. Its bigger limit is the login wall: logged out, a profile shows five posts, a post shows a few replies, and search shows nothing.
  • In the September 16, 2026 benchmark, 14 of 16 scraping APIs returned x.com on all five attempts. X ranked 86th of 100 sites by difficulty. ScraperAPI and Scrapingdog passed none.
  • That benchmark row is easier than the job. The pass check is the text "NASA", which any page about the account contains. It does not count posts.
  • Through String, the profile page had 5 posts on each of four fetches, and a post page had the post and 3 of its 192 replies. Neither needed JavaScript rendering.
  • Both billed as a standard fetch: $0.30 per 1,000 on Starter and $0.20 on Growth.

What people who scrape X run into

Everyone who scraped Twitter remembers the date. On June 30, 2023, the maintainer of snscrape, the most used open-source Twitter scraper, wrote that "all Twitter scrapes are failing since sometime in the past hour" and tied it to "Twitter as a whole getting locked behind a login wall" (snscrape issue #996, June 30, 2023). Earlier that day a new user had asked in the same tracker whether the change would matter: "I'm new to scraping so not sure if this affects scrapers" (snscrape issue #994, June 30, 2023). It did. Search-based scraping never came back for logged-out clients.

What came back, partly, is the public profile and post page. X now serves both to logged-out visitors with real content in the HTML, which is what the scripts below read. The official route is the X API, which its docs describe as "pay-per-usage pricing" with credits bought up front. If you need search, full timelines or follower lists, the API is the route that includes them.

What a logged-out request to X gets

We requested NASA's profile, one NASA post and a search URL on September 30, 2026. The plain request ran from a home connection outside the US; the String requests asked for a US exit. The profile is the benchmark's X URL.

What we tried, September 30, 2026 Profile x.com/NASA Post x.com/NASA/status/... Search x.com/search?q=nasa
Plain requests from a home IP HTTP 200, 181 KB, bio and counts, 5 posts Not tried Not tried
String fetch HTTP 200, 181 KB, bio and counts, 5 posts (4 of 4 tries) HTTP 200, 129 KB, the post and 3 replies HTTP 200, 17 KB, "X - The Everything App", no results
String fetch, executeJS: true HTTP 200, 129 KB, bio and counts, 5 posts Not tried Not tried

Two results matter. First, a plain request from a home connection outside the US got the same five posts as String did. X is not hard to reach from a home IP at low volume; we did not test it from a cloud server, where most scrapers run. Second, search is gone for logged-out clients. The 17 KB page is the app shell with nothing in it, so a search scraper against x.com needs a logged-in session or the API. Rendering with executeJS returned the same five posts and billed as a browser fetch, which costs five times as much, so leave it off.

Ways to get X data, compared

Method Profile and latest posts Search, full timelines What you maintain
Plain requests Profile and 5 posts from a home IP in our test; untested from a server No Parsers for HTML X changes
snscrape No, failing since the June 2023 login wall No Nothing that works
Headless browser, logged in Yes Yes Accounts, sessions, and the risk of those accounts being suspended
The X API Yes Yes Credits and X's developer terms
A fetch API such as the String Web Access API Yes No Your parser, which is the HTML read below

What 16 scraping APIs did with x.com

The Web Data Frontier Benchmark sends the same 100 URLs to 16 scraping APIs, five attempts each. The X URL is NASA's profile. The harness and targets are open source, and so is the raw result file for the September 16 run.

API x.com, 5 attempts Average time on x.com Overall success, 100 sites
String 100% 1.6 s 97.0%
Scrapfly 100% 13.1 s 86.2%
ScraperAPI 0% none passed 84.0%
Firecrawl 100% 6.2 s 80.2%
Apify 100% 3.5 s 77.4%
Bright Data 100% 3.5 s 74.6%
ScrapingBee 100% 2.0 s 73.0%
Context.dev 100% 4.6 s 72.0%
Oxylabs 100% 3.0 s 69.0%
Nimble 100% 2.8 s 68.6%
Zyte 100% 1.5 s 68.0%
Decodo 100% 32.2 s 50.6%
Scrapingdog 0% none passed 45.6%
Browserbase 100% 2.2 s 41.4%
ZenRows 100% 5.3 s 41.2%
ScrapingAnt 100% 8.4 s 36.4%

Fourteen of 16 APIs returned x.com on all five attempts: 70 passes out of 80 requests. The average API passed 87.5% of attempts on X, against 66.6% across all 100 sites. X ranked 86th of 100 by difficulty.

Do not read that as "X is easy to scrape." The benchmark counts a pass when the response is a 2xx and the body contains "NASA". The profile page we got contains "NASA" 135 times, so the check is met by any page about the account, with or without posts. This row measures whether an API gets past X's bot defense to the profile page. It does not measure what came with it, and it says nothing about search, which no logged-out client gets. Test any tool, including ours, by counting posts.

How to scrape an X profile with Python

Install requests, set your key, and pass a handle. It reads the follower count and each post: its ID, date, text, replies, reposts, likes and views. Python 3.10 or later.

"""Scrape a public X (Twitter) profile and its latest posts with the String Web Access API.

    pip install requests
    export STRING_API_KEY=...
    python x_profile.py NASA
"""
import html
import os
import re
import sys
from datetime import datetime, timezone

import requests

API = "https://request.usestring.ai/v1/fetch"


def fetch(url: str) -> str:
    response = requests.post(
        API,
        headers={"Authorization": f"Bearer {os.environ['STRING_API_KEY']}"},
        json={"url": url, "countryCode": "US"},
        timeout=90,
    )
    response.raise_for_status()
    page = response.json()
    if page["statusCode"] != 200:
        raise RuntimeError(f"x.com returned {page['statusCode']} for {url}")
    return page["data"]


def text(fragment: str) -> str:
    return html.unescape(re.sub(r"\s+", " ", re.sub(r"<[^>]+>", "", fragment))).strip()


def count_after(label: str, fragment: str) -> str | None:
    m = re.search(rf'aria-label="{label}".*?>([\d.,]+[KMB]?)<', fragment, re.S)
    return m.group(1) if m else None


def profile(handle: str) -> dict:
    page = fetch(f"https://x.com/{handle}")
    if f"(@{handle}) / X" not in page:
        raise RuntimeError("no profile title in the page: treat this as blocked or a login wall")
    followers = re.search(r'>([\d.,]+[KMB]?)</div><div[^>]*>Followers<', page)
    posts = []
    for article in re.findall(r"<article.*?</article>", page, re.S):
        status = re.search(rf'href="/{handle}/status/(\d+)"', article, re.I)
        ts = re.search(r'"timestamp":(\d+)', article)
        body = re.search(r'<div dir="auto"[^>]*>(.*?)</div>', article, re.S)
        if not status:
            continue  # a repost or a card from another account
        posts.append({
            "id": status.group(1),
            "date": datetime.fromtimestamp(int(ts.group(1)) / 1000, timezone.utc).date().isoformat() if ts else None,
            "text": text(body.group(1)) if body else "",
            "replies": count_after("Reply", article),
            "reposts": count_after("Repost", article),
            "likes": count_after("Like", article),
            "views": count_after("View count", article),
        })
    return {"handle": handle, "followers": followers.group(1) if followers else None, "posts": posts}


if __name__ == "__main__":
    data = profile(sys.argv[1] if len(sys.argv) > 1 else "NASA")
    print(f"@{data['handle']}: {data['followers']} followers, {len(data['posts'])} posts in the page")
    for p in data["posts"]:
        print(f"{p['date']}  {p['likes']:>5} likes {p['views']:>6} views  {p['text'][:55]}")

What it printed on September 30, 2026, in 2.1 seconds:

@NASA: 92.3M followers, 5 posts in the page
2026-09-28     3K likes   652K views  How do we respond when we detect an asteroid that could
2026-09-29   1.1K likes   385K views  We're ready for Crew-13 to launch to the @Space_Station
2026-09-28   2.2K likes   681K views  On Monday, NASA and @BoeingSpace provided an update on 
2026-09-28   1.2K likes   818K views  LIVE: NASA and @BoeingSpace leaders give updates on the
2026-09-25   4.4K likes   966K views  With @BoeingSpace, we’ll provide an update on Starliner

The first post is older than the second because it is pinned. The counts are X's rounded display values, such as "3K" and "651K", not exact numbers. The five posts are all a logged-out visitor gets. To follow an account over time, run the script on a schedule and keep the post IDs you have not seen before.

How to scrape an X post with TypeScript

No dependencies. Node 22.6 or later runs TypeScript directly (we ran it on Node 24), and so does Bun.

// Scrape one public X (Twitter) post and the replies X shows logged-out, with the String Web Access API.
//   export STRING_API_KEY=...
//   node x_post.ts <status url>      (Node 22.6+; or: bun x_post.ts <url>)

async function fetchPage(url: string): Promise<string> {
  const response = await fetch("https://request.usestring.ai/v1/fetch", {
    method: "POST",
    headers: { Authorization: `Bearer ${process.env.STRING_API_KEY}`, "Content-Type": "application/json" },
    body: JSON.stringify({ url, countryCode: "US" }),
  });
  if (!response.ok) throw new Error(`String returned ${response.status}: ${await response.text()}`);
  const page = (await response.json()) as { statusCode: number; data: string };
  if (page.statusCode !== 200) throw new Error(`x.com returned ${page.statusCode}`);
  return page.data;
}

const decode = (s: string) =>
  s.replace(/&quot;/g, '"').replace(/&#x27;|&#39;/g, "'").replace(/&amp;/g, "&").replace(/&lt;/g, "<").replace(/&gt;/g, ">");
const plain = (s: string) => decode(s.replace(/<[^>]+>/g, "")).replace(/\s+/g, " ").trim();
const countAfter = (label: string, s: string) =>
  s.match(new RegExp(`aria-label="${label}"[\\s\\S]*?>([\\d.,]+[KMB]?)<`))?.[1] ?? "n/a";

const url = process.argv[2] ?? "https://x.com/NASA/status/2104593672624886205";
const [, handle, id] = url.match(/x\.com\/([^/]+)\/status\/(\d+)/) ?? [];
if (!id) throw new Error("pass a https://x.com/<handle>/status/<id> URL");

const html = await fetchPage(url);
const articles = html.match(/<article[\s\S]*?<\/article>/g) ?? [];
const post = articles.find((a) => a.includes(`/status/${id}"`));
if (!post) throw new Error("HTTP 200 but the post is not in the page: treat this as a login wall or a deleted post");

const ts = Number(post.match(/"timestamp":(\d+)/)?.[1] ?? html.match(/"timestamp":(\d+)/)?.[1]);
const body = post.match(/<div dir="auto"[^>]*>([\s\S]*?)<\/div>/)?.[1] ?? "";
// On a post page the view count is a "651.7K Views" line, not a button.
const views = post.match(/>([\d.,]+[KMB]?)<\/div><div[^>]*>Views</)?.[1] ?? "not shown";
const replies = articles.filter((a) => !a.includes(`/status/${id}"`)).length;

console.log(`@${handle} ${id} | ${ts ? new Date(ts).toISOString().slice(0, 10) : "no date"}`);
console.log(plain(body).slice(0, 100));
console.log(`replies ${countAfter("Reply", post)} | reposts ${countAfter("Repost", post)} | likes ${countAfter("Like", post)} | views ${views}`);
console.log(`${replies} replies included in the logged-out page`);

We ran it on September 30, 2026, five times across Node 24 and Bun. Runs took 1.7 to 9.6 seconds. The last run printed:

@NASA 2104593672624886205 | 2026-09-28
How do we respond when we detect an asteroid that could pose a threat to Earth? Follow the story beh
replies 192 | reposts 494 | likes 3K | views 652K
3 replies included in the logged-out page

The post had 192 replies and the page carried 3 of them. That is the login wall again. Each reply in the page is its own <article> with the same fields as the post, so the script can be extended to read them; we only count them here.

The failure mode: a page that is real and incomplete

X's limits look like data. Every page we fetched came back HTTP 200 with the right title, and each one held less than the account has:

  • The profile holds five posts. NASA has 74.3K posts; the page carried 5, and the first one is pinned, so it is not the newest. A scraper that treats the page as the timeline will miss everything past the fifth post.
  • The post holds three replies. The post had 192 replies; the page carried 3.
  • Search holds nothing. The search URL returned a 17 KB app shell with no posts and the generic title "X - The Everything App".

Three checks catch the ways this goes wrong:

  • Check the title. The Python script raises an error if "(@handle) / X" is missing, which catches the app shell X sends for search and for pages it will not show logged out.
  • Check the post ID. The TypeScript script looks for the exact status ID in the page. A deleted post or a protected account gives a page without it.
  • Count the posts. A profile page for an active account with zero <article> elements is a partial page. Refetch it before you record that the account went quiet.

The other risk is the parser. X's markup uses generated class names that change. Both scripts read <article>, aria-label and the /status/ link rather than class names, and they will still break when X changes its page. When a count is missing, the Python script returns None and the TypeScript prints "n/a" rather than a zero, so a broken parser does not look like a post nobody liked.

We did not check X pages for location. Nothing in these pages told us which country the request came from.

Instagram and TikTok

The bank question this page answers asks about X, Instagram and TikTok together. Here is what the benchmark and one fetch each showed for the other two. They will get their own pages.

Instagram (instagram.com/nasa/) TikTok (tiktok.com/@nba)
APIs that passed 5 of 5, of 16 11 10
Passes out of 80 attempts 59 64
Difficulty rank, of 100 59th 72nd
APIs that passed none ScraperAPI, Firecrawl, Scrapingdog, ZenRows Firecrawl, ZenRows
String, 5 attempts and average time 100%, 2.3 s 100%, 1.1 s
One String fetch, September 30, 2026 739 KB; the page description carried "104M Followers, 93 Following, 4,937 Posts" 374 KB; JSON in __UNIVERSAL_DATA_FOR_REHYDRATION__ with a follower count of 27,200,000 and a video count of 23,800, and no video list
Billed as Premium fetch Browser fetch

The pattern is close to X's. Logged out, each site gave us the profile's counts. The TikTok page carried no video list, and we did not check Instagram's page for posts. On both, as on X, the benchmark's pass check is the account name, which a page with counts and no posts satisfies. String also publishes a TikTok profile Actor on Apify that returns exact follower counts rather than rounded ones.

What it costs

From the billing headers we logged on September 30, 2026, and String's pricing:

  • X profile or post page. Every x.com request we logged without rendering billed as a standard fetch: $0.30 per 1,000 on Starter ($20 a month) and $0.20 on Growth ($100 a month). Checking 1,000 accounts once a day is 30,000 requests a month: $9 on Starter or $6 on Growth.
  • With executeJS on. Billed as a browser fetch: $1.50 per 1,000 on Starter and $1.00 on Growth. It returned the same five posts, so you do not need it for X.
  • Instagram and TikTok profiles. The Instagram fetch billed as a premium fetch, $3.00 per 1,000 on Starter and $2.00 on Growth. The TikTok fetch billed as a browser fetch, $1.50 and $1.00.

A request that fails is not billed. If you need search or full timelines, price the X API instead, because a logged-out page cannot give you either.

FAQ

How to scrape Twitter (X) in 2026?

Fetch the public profile or post page and read the posts from the <article> elements. Logged out, X shows a profile's bio, follower count and five latest posts, and a post with its first few replies. The Python script on this page returned NASA's profile and five posts in 2.1 seconds on September 30, 2026.

Can you scrape X without logging in?

Yes, for public profiles and single posts. In our test on September 30, 2026, a logged-out profile page carried five posts and a post page carried the post and three of its 192 replies. Search returned an empty app shell. Anything beyond that needs a logged-in session or the X API.

Why does my X scraper only get five posts?

That is all X shows a logged-out visitor on a profile page. In our test on September 30, 2026, NASA's profile page carried 5 of its 74.3K posts, and one of the five was pinned. To build a history, run the scraper on a schedule and keep new post IDs, or use the X API for full timelines.

Does snscrape still work for Twitter?

No. Its maintainer wrote on June 30, 2023 that "all Twitter scrapes are failing", after X put the site behind a login wall, and the issue was never closed with a fix. In our test, search is still an empty page for logged-out clients.

Does X have an official API?

Yes. The X API uses what its docs call "pay-per-usage pricing": you buy credits and each call draws them down. It is the supported route to search, timelines and follower lists, which a logged-out page does not show.

How do I scrape Instagram and TikTok?

The same way as X: fetch the public profile page and read what the site puts in it for a logged-out visitor. In one fetch each on September 30, 2026, Instagram's page description carried follower, following and post counts, and TikTok's page carried a JSON block with exact follower and video counts but no video list.

Is it legal to scrape social media?

It depends on what you collect and where you are. Public, logged-out pages are a different case from content behind a login, and personal data brings privacy law with it. Read our pages on whether web scraping is legal and the hiQ v. LinkedIn ruling, and ask a lawyer about your use.

Sources

Cheers,
String team

Get your API key →Explore the Web Access API
© 2026 StringEU and UK GDPR Article 27 representative — appointment verified by EuverifyBuilt in New York City 🗽 🍎