Claude does, and ChatGPT does only in part. Anthropic says all three of its bots honor robots.txt, including Claude-User, the agent that fetches a page when you ask Claude about it. OpenAI's crawlers, GPTBot and OAI-SearchBot, honor it too. But OpenAI's own documentation says that for ChatGPT-User, the agent that fetches a page when you ask ChatGPT a question, "robots.txt rules may not apply." Sites are acting on this: on October 6, 2026 we read the robots.txt files of the 100 sites in String's benchmark, and 23 of the 98 we could read disallow ChatGPT-User on the page we test, while 28 disallow Claude-User.
Each company runs separate agents for training, search and user-requested fetches, and each one reads its own robots.txt line.
| Company | Agent | What it does | Does it follow robots.txt? |
|---|---|---|---|
| OpenAI | GPTBot | Collects pages that may train OpenAI's models | Yes. Disallow means "do not use for training" |
| OpenAI | OAI-SearchBot | Indexes pages for ChatGPT search results | Yes. Disallow removes the site from ChatGPT search answers |
| OpenAI | ChatGPT-User | Fetches a page when a user asks ChatGPT or a custom GPT | "Robots.txt rules may not apply," per OpenAI |
| Anthropic | ClaudeBot | Collects pages that may train Claude | Yes |
| Anthropic | Claude-SearchBot | Indexes pages to improve Claude's search results | Yes |
| Anthropic | Claude-User | Fetches a page when a user asks Claude | Yes. Disallow stops Claude retrieving the page for a user |
| Perplexity | PerplexityBot | Indexes pages for Perplexity search | Yes |
| Perplexity | Perplexity-User | Fetches a page when a user asks Perplexity | "Generally ignores robots.txt rules," per Perplexity |
Sources: OpenAI's crawler overview, Anthropic's crawler article (April 7, 2026), and Perplexity's crawler page. OpenAI gives its reason in one line: "Because these actions are initiated by a user, robots.txt rules may not apply." Anthropic takes the other view and adds that its bots "respect anti-circumvention technologies," such as not trying to get past CAPTCHAs on sites they crawl.
We fetched /robots.txt from each of the 100 sites in String's Web Data Frontier Benchmark on October 6, 2026, and read 98 (Stack Overflow and eMAG refused both a plain client and a fetch API). For each AI agent we asked two things: does the file name the agent, and does it allow that agent to fetch the exact page the benchmark tests? We followed RFC 9309: a group that names the agent wins over the * group, and the longest matching rule wins.
| Agent | Files that name it | Files that disallow the benchmark page for it |
|---|---|---|
| GPTBot | 34 | 35 |
| OAI-SearchBot | 30 | 22 |
| ChatGPT-User | 30 | 23 |
| ClaudeBot | 28 | 34 |
| Claude-SearchBot | 17 | 27 |
| Claude-User | 18 | 28 |
| PerplexityBot | 26 | 29 |
| Perplexity-User | 14 | 29 |
| Google-Extended | 25 | 34 |
| CCBot | 21 | 38 |
What the table shows:
* group.* group, which also covers every other bot. Google's own results page is disallowed for everyone.A disallow for Claude-User is a rule Anthropic says it follows. A disallow for ChatGPT-User is, in OpenAI's words, a rule that may not apply.
Policy and behavior are measured in two places, and they do not fully agree.
Agent features that drive a full browser are a separate case. OpenAI's cloud browser for ChatGPT signs each request with Web Bot Auth (HTTP Message Signatures, RFC 9421) and a Signature-Agent header set to https://chatgpt.com. OpenAI's guidance for site owners is about allowlisting that traffic in Akamai, Cloudflare, HUMAN or Vercel, not about robots.txt (OpenAI Help Center). If you want to control a signed agent, do it at your CDN or firewall.
Name every agent you mean, because each one obeys only the group that names it, or * when none does. To stay in AI search answers but opt out of training:
User-agent: OAI-SearchBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /
User-agent: GPTBot
User-agent: ClaudeBot
User-agent: CCBot
User-agent: Google-Extended
Disallow: /
Add ChatGPT-User, Claude-User and Perplexity-User to the disallow group only if you do not want assistants to read a page a user points them at. Claude-User will follow it. ChatGPT-User and Perplexity-User, by their operators' own documentation, may not. For a page that must stay private, use a login or a CDN rule, not robots.txt. Our page on whether robots.txt is legally binding covers what the file does and does not do in law.
For reference, usestring.ai allows every crawler and states its preferences with a Content-Signal line: search=yes, ai-input=yes, ai-train=no.
If your product fetches pages for users, decide your own rule and enforce it in code, so the behavior is yours and visible in your logs. This Python function reads the site's robots.txt for your agent's token and applies RFC 9309 matching.
import re
import sys
from urllib.parse import urlparse
import requests
def groups(text: str) -> dict:
out, agents, rules, after_agent = {}, [], [], False
for line in text.splitlines():
line = line.split("#", 1)[0].strip()
if ":" not in line:
continue
field, value = (x.strip() for x in line.split(":", 1))
field = field.lower()
if field == "user-agent":
if not after_agent:
agents, rules = [], []
agents.append(value.lower())
out.setdefault(value.lower(), rules)
after_agent = True
elif field in ("allow", "disallow"):
after_agent = False
rules.append((field == "allow", value))
else:
after_agent = False
return out
def matches(pattern: str, path: str) -> bool:
if not pattern:
return False
rx = re.escape(pattern).replace(r"\*", ".*")
if rx.endswith(r"\$"):
rx = rx[:-2] + "$"
return re.match(rx, path) is not None
def may_fetch(token: str, url: str) -> bool:
u = urlparse(url)
r = requests.get(f"{u.scheme}://{u.netloc}/robots.txt", timeout=15,
headers={"User-Agent": f"{token}/1.0"})
if r.status_code in (404, 410):
return True # no file: everything is allowed
r.raise_for_status() # 401/403/5xx: treat as "do not fetch" and look by hand
g = groups(r.text)
rules = g.get(token.lower(), g.get("*", []))
path = (u.path or "/") + (f"?{u.query}" if u.query else "")
best = max(((len(p), allow) for allow, p in rules if matches(p, path)), default=None)
return True if best is None else best[1]
if __name__ == "__main__":
token, urls = sys.argv[1], sys.argv[2:]
for url in urls:
try:
print(token, url, "allowed" if may_fetch(token, url) else "disallowed")
except requests.HTTPError as e:
print(token, url, f"robots.txt returned {e.response.status_code}: check by hand")
We ran it on October 6, 2026. As Claude-User, an Amazon product page and a New York Times section page came back disallowed and a Wikipedia article came back allowed. As a new agent name with no rules of its own, all three were allowed, and a Google results page was disallowed under every name. Use your own token. Amazon and Wikipedia refused our robots.txt request when it claimed to be ChatGPT-User.
String is the most reliable web scraping API, bypassing anti-bot, and you only pay when content comes back. For an agent that wants to follow site rules, three parts of it matter:
403 with a reason field, so your agent can branch on it (access control).x-status-code header carries the site's response, so your agent sees a 403 or 429 from the site and can stop or slow down.Pair it with the may_fetch check above: your code decides the robots.txt policy, and String returns the page when the answer is yes.
String's Web Access API gives agents search, fetch to clean Markdown, and a hosted MCP server.
Put the may_fetch check in front of your fetch call and log its answer, and your agent's robots.txt policy is something you can show a site owner.
Claude does: Anthropic says all three of its bots, including Claude-User for user-requested fetches, honor robots.txt. ChatGPT's crawlers GPTBot and OAI-SearchBot honor it, but OpenAI says robots.txt rules may not apply to ChatGPT-User, which fetches pages when a user asks.
Not reliably. OpenAI's documentation says that because ChatGPT-User acts on a user's request, "robots.txt rules may not apply." TollBit data for the first half of 2026 found it reaching disallowed pages on nearly half of the European sites that listed it, though one controlled test in June 2026 saw it comply.
Yes, by Anthropic's account. Disallowing Claude-User stops Claude retrieving your pages when a user asks about them. On October 6, 2026, 28 of 98 benchmark sites we checked disallowed Claude-User on the page we test.
Disallow GPTBot to opt out of training and OAI-SearchBot to leave ChatGPT search. Add ChatGPT-User to cover user-requested fetches, knowing OpenAI says that rule may not apply. For a hard block, use your CDN or firewall with OpenAI's published IP ranges.
Yes. GPTBot controls training only. Search visibility is controlled by OAI-SearchBot, so a site can block GPTBot and still be cited in ChatGPT search answers. Of the 98 benchmark sites we checked, 35 disallow GPTBot on the page we test and 22 disallow OAI-SearchBot.
No. ClaudeBot is the training crawler. Claude-User fetches pages for users and Claude-SearchBot indexes for search, and each reads its own robots.txt line. Disallow all three if you want Claude to stay off the site.
PerplexityBot, its search crawler, does. Perplexity says Perplexity-User, which fetches pages when a user asks, "generally ignores robots.txt rules" because a user requested the fetch.
Not in itself in most places. Robots.txt is a convention, not a contract or an access control, though ignoring it can matter in a dispute about terms of service or unauthorized access. Our page on whether robots.txt legally prevents scraping covers the cases; this is not legal advice.
/robots.txt from the 100 sites in String's Web Data Frontier Benchmark, 98 read, evaluated per agent against the benchmark page under RFC 9309 rules. Target list from official_results/benchmark-2026-09-16T01-03-47-074Z.json in the public harness.