NewLaunching String Web Access APIRead the manifesto →
← Answers

Do ChatGPT and Claude respect robots.txt?

String team · Updated October 6, 2026

Claude does, and ChatGPT does only in part. Anthropic says all three of its bots honor robots.txt, including Claude-User, the agent that fetches a page when you ask Claude about it. OpenAI's crawlers, GPTBot and OAI-SearchBot, honor it too. But OpenAI's own documentation says that for ChatGPT-User, the agent that fetches a page when you ask ChatGPT a question, "robots.txt rules may not apply." Sites are acting on this: on October 6, 2026 we read the robots.txt files of the 100 sites in String's benchmark, and 23 of the 98 we could read disallow ChatGPT-User on the page we test, while 28 disallow Claude-User.

Which bot does what

Each company runs separate agents for training, search and user-requested fetches, and each one reads its own robots.txt line.

Company Agent What it does Does it follow robots.txt?
OpenAI GPTBot Collects pages that may train OpenAI's models Yes. Disallow means "do not use for training"
OpenAI OAI-SearchBot Indexes pages for ChatGPT search results Yes. Disallow removes the site from ChatGPT search answers
OpenAI ChatGPT-User Fetches a page when a user asks ChatGPT or a custom GPT "Robots.txt rules may not apply," per OpenAI
Anthropic ClaudeBot Collects pages that may train Claude Yes
Anthropic Claude-SearchBot Indexes pages to improve Claude's search results Yes
Anthropic Claude-User Fetches a page when a user asks Claude Yes. Disallow stops Claude retrieving the page for a user
Perplexity PerplexityBot Indexes pages for Perplexity search Yes
Perplexity Perplexity-User Fetches a page when a user asks Perplexity "Generally ignores robots.txt rules," per Perplexity

Sources: OpenAI's crawler overview, Anthropic's crawler article (April 7, 2026), and Perplexity's crawler page. OpenAI gives its reason in one line: "Because these actions are initiated by a user, robots.txt rules may not apply." Anthropic takes the other view and adds that its bots "respect anti-circumvention technologies," such as not trying to get past CAPTCHAs on sites they crawl.

What 98 robots.txt files say about these agents

We fetched /robots.txt from each of the 100 sites in String's Web Data Frontier Benchmark on October 6, 2026, and read 98 (Stack Overflow and eMAG refused both a plain client and a fetch API). For each AI agent we asked two things: does the file name the agent, and does it allow that agent to fetch the exact page the benchmark tests? We followed RFC 9309: a group that names the agent wins over the * group, and the longest matching rule wins.

Agent Files that name it Files that disallow the benchmark page for it
GPTBot 34 35
OAI-SearchBot 30 22
ChatGPT-User 30 23
ClaudeBot 28 34
Claude-SearchBot 17 27
Claude-User 18 28
PerplexityBot 26 29
Perplexity-User 14 29
Google-Extended 25 34
CCBot 21 38

What the table shows:

  • 43 of 98 sites name at least one AI agent. The other 55 handle AI traffic, if at all, through the * group.
  • 22 sites disallow both user-requested fetchers, ChatGPT-User and Claude-User, on the tested page. Amazon, The New York Times, Yelp, LinkedIn and Reddit are among them.
  • 13 sites block GPTBot but allow ChatGPT-User. These sites say no to training and yes to a user asking about a page.
  • 20 files name ChatGPT-User only to allow it. Claude-User is named and allowed in 11.
  • Many blocks are not aimed at AI at all. For Claude-User, 21 of the 28 disallows come from the * group, which also covers every other bot. Google's own results page is disallowed for everyone.

A disallow for Claude-User is a rule Anthropic says it follows. A disallow for ChatGPT-User is, in OpenAI's words, a rule that may not apply.

What happens in practice

Policy and behavior are measured in two places, and they do not fully agree.

  • A controlled test found all three compliant. Search Engine World gave 12 assistants their own secret URL plus a twin URL that robots.txt disallowed, then read the server logs. Only ChatGPT, Claude and Perplexity left the disallowed twin alone; the other nine fetched it (Search Engine World, June 2026). That was one site and one round.
  • Large-scale logs show ChatGPT-User reaching disallowed pages. Search Engine Journal, reporting TollBit's State of the Bots data for the first half of 2026, found that ChatGPT-User, Bytespider and Youbot each accessed disallowed pages on nearly half of the European sites that had listed them, with ChatGPT-User on the most sites (Search Engine Journal, August 14, 2026).
  • Some sites check the name. From our own connection, a request for Amazon's robots.txt that called itself ChatGPT-User got a 503, and the same request to Wikipedia got a 403. Called Claude-User or a made-up agent name, both returned 200. A user-agent string is a claim, and some large sites check it.

What about agent mode and AI browsers

Agent features that drive a full browser are a separate case. OpenAI's cloud browser for ChatGPT signs each request with Web Bot Auth (HTTP Message Signatures, RFC 9421) and a Signature-Agent header set to https://chatgpt.com. OpenAI's guidance for site owners is about allowlisting that traffic in Akamai, Cloudflare, HUMAN or Vercel, not about robots.txt (OpenAI Help Center). If you want to control a signed agent, do it at your CDN or firewall.

If you run a site

Name every agent you mean, because each one obeys only the group that names it, or * when none does. To stay in AI search answers but opt out of training:

User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /

User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
Disallow: /

Add ChatGPT-User, Claude-User and Perplexity-User to the disallow group only if you do not want assistants to read a page a user points them at. Claude-User will follow it. ChatGPT-User and Perplexity-User, by their operators' own documentation, may not. For a page that must stay private, use a login or a CDN rule, not robots.txt. Our page on whether robots.txt is legally binding covers what the file does and does not do in law.

For reference, usestring.ai allows every crawler and states its preferences with a Content-Signal line: search=yes, ai-input=yes, ai-train=no.

If you build an agent: check robots.txt yourself

If your product fetches pages for users, decide your own rule and enforce it in code, so the behavior is yours and visible in your logs. This Python function reads the site's robots.txt for your agent's token and applies RFC 9309 matching.

import re
import sys
from urllib.parse import urlparse

import requests


def groups(text: str) -> dict:
    out, agents, rules, after_agent = {}, [], [], False
    for line in text.splitlines():
        line = line.split("#", 1)[0].strip()
        if ":" not in line:
            continue
        field, value = (x.strip() for x in line.split(":", 1))
        field = field.lower()
        if field == "user-agent":
            if not after_agent:
                agents, rules = [], []
            agents.append(value.lower())
            out.setdefault(value.lower(), rules)
            after_agent = True
        elif field in ("allow", "disallow"):
            after_agent = False
            rules.append((field == "allow", value))
        else:
            after_agent = False
    return out


def matches(pattern: str, path: str) -> bool:
    if not pattern:
        return False
    rx = re.escape(pattern).replace(r"\*", ".*")
    if rx.endswith(r"\$"):
        rx = rx[:-2] + "$"
    return re.match(rx, path) is not None


def may_fetch(token: str, url: str) -> bool:
    u = urlparse(url)
    r = requests.get(f"{u.scheme}://{u.netloc}/robots.txt", timeout=15,
                     headers={"User-Agent": f"{token}/1.0"})
    if r.status_code in (404, 410):
        return True                       # no file: everything is allowed
    r.raise_for_status()                  # 401/403/5xx: treat as "do not fetch" and look by hand
    g = groups(r.text)
    rules = g.get(token.lower(), g.get("*", []))
    path = (u.path or "/") + (f"?{u.query}" if u.query else "")
    best = max(((len(p), allow) for allow, p in rules if matches(p, path)), default=None)
    return True if best is None else best[1]


if __name__ == "__main__":
    token, urls = sys.argv[1], sys.argv[2:]
    for url in urls:
        try:
            print(token, url, "allowed" if may_fetch(token, url) else "disallowed")
        except requests.HTTPError as e:
            print(token, url, f"robots.txt returned {e.response.status_code}: check by hand")

We ran it on October 6, 2026. As Claude-User, an Amazon product page and a New York Times section page came back disallowed and a Wikipedia article came back allowed. As a new agent name with no rules of its own, all three were allowed, and a Google results page was disallowed under every name. Use your own token. Amazon and Wikipedia refused our robots.txt request when it claimed to be ChatGPT-User.

How String helps an agent read the web responsibly

String is the most reliable web scraping API, bypassing anti-bot, and you only pay when content comes back. For an agent that wants to follow site rules, three parts of it matter:

  • It reads robots.txt files a plain client cannot. In our October 6, 2026 survey, 21 of the 98 robots.txt files we read came back only through String; a plain request got a challenge or a block. An agent cannot follow a file it cannot read.
  • It gates sensitive destinations. Some sites, such as social networks, search engines, forums and job sites, need approval before String fetches them. An unapproved domain returns 403 with a reason field, so your agent can branch on it (access control).
  • It re-checks every navigation. When an agent drives a page with browser actions, String checks each click, redirect and script navigation against the same rules, not only the first URL.
  • It reports the site's own status. The x-status-code header carries the site's response, so your agent sees a 403 or 429 from the site and can stop or slow down.

Pair it with the may_fetch check above: your code decides the robots.txt policy, and String returns the page when the answer is yes.

Where String fits

String's Web Access API gives agents search, fetch to clean Markdown, and a hosted MCP server.

  • For: teams building agents and data pipelines that read public pages and want a fetch that returns content on bot-protected sites. In the September 16, 2026 benchmark run, String returned content on 97.0% of 500 requests to 100 bot-protected pages, the highest of 16 APIs.
  • Not for: content behind a login you do not own.
  • Price: Starter is $20 a month; a standard fetch is $0.30 per 1,000 successful requests, and you pay only when content comes back. The first 5,000 standard requests are free (pricing).
  • Integrations: a REST API, and a hosted MCP server for MCP clients such as Claude.
  • Limits: 60 requests a second per account with a 3,600-request burst; concurrency is not capped.

Put the may_fetch check in front of your fetch call and log its answer, and your agent's robots.txt policy is something you can show a site owner.

FAQ

Do ChatGPT and Claude respect robots.txt when browsing?

Claude does: Anthropic says all three of its bots, including Claude-User for user-requested fetches, honor robots.txt. ChatGPT's crawlers GPTBot and OAI-SearchBot honor it, but OpenAI says robots.txt rules may not apply to ChatGPT-User, which fetches pages when a user asks.

Does ChatGPT-User respect robots.txt?

Not reliably. OpenAI's documentation says that because ChatGPT-User acts on a user's request, "robots.txt rules may not apply." TollBit data for the first half of 2026 found it reaching disallowed pages on nearly half of the European sites that listed it, though one controlled test in June 2026 saw it comply.

Does Claude-User respect robots.txt?

Yes, by Anthropic's account. Disallowing Claude-User stops Claude retrieving your pages when a user asks about them. On October 6, 2026, 28 of 98 benchmark sites we checked disallowed Claude-User on the page we test.

How do I block ChatGPT from my website?

Disallow GPTBot to opt out of training and OAI-SearchBot to leave ChatGPT search. Add ChatGPT-User to cover user-requested fetches, knowing OpenAI says that rule may not apply. For a hard block, use your CDN or firewall with OpenAI's published IP ranges.

If I block GPTBot, will my site still appear in ChatGPT?

Yes. GPTBot controls training only. Search visibility is controlled by OAI-SearchBot, so a site can block GPTBot and still be cited in ChatGPT search answers. Of the 98 benchmark sites we checked, 35 disallow GPTBot on the page we test and 22 disallow OAI-SearchBot.

Does blocking ClaudeBot stop Claude from reading my page?

No. ClaudeBot is the training crawler. Claude-User fetches pages for users and Claude-SearchBot indexes for search, and each reads its own robots.txt line. Disallow all three if you want Claude to stay off the site.

Does Perplexity respect robots.txt?

PerplexityBot, its search crawler, does. Perplexity says Perplexity-User, which fetches pages when a user asks, "generally ignores robots.txt rules" because a user requested the fetch.

Is ignoring robots.txt illegal?

Not in itself in most places. Robots.txt is a convention, not a contract or an access control, though ignoring it can matter in a dispute about terms of service or unauthorized access. Our page on whether robots.txt legally prevents scraping covers the cases; this is not legal advice.

Sources

Get your API key →Explore the Web Access API
© 2026 StringEU and UK GDPR Article 27 representative — appointment verified by EuverifyBuilt in New York City 🗽 🍎