How Google parses robots.txt — and where our validator was wrong
There is a right answer, and it is runnable
robots.txt has a standard, RFC 9309, written in 2022 by engineers at Google. Google also publishes the C++ parser it uses in production (github.com/google/robotstxt), with a command-line tool that takes a file, a crawler name and a path, and answers allowed or not. That makes it something rare in SEO: a claim you can check rather than argue about.
Our validator was written from the RFC and tested against our own reading of it. On 23 September 2026 we wrote 61 cases by hand — the shapes we had seen in client files — and asked both parsers. Seventeen verdicts differed, and in every one of them Google was reading the standard correctly and we were not. The fixture is now 72 cases, each answer recorded from Google's binary, and the test fails the day our parser stops agreeing. We also generated 21,000 random files and 30 real ones and found no further disagreement.
Four rules carry nearly every verdict
1. A crawler that is named stops reading the wildcard group
A crawler looks for groups whose User-agent names it. If any do, those are the only groups it reads; the User-agent: * group is no longer about it. Only a crawler nothing names falls back to *. This is the rule behind the most common silent change on a site: adding User-agent: GPTBot with one line under it, and thereby releasing GPTBot from every rule the site had written for everyone.
User-agent: *Allow: /Disallow: /hooks/Disallow: /app/Disallow: /api/Disallow: /sso/# … 24 Disallow lines in all …User-agent: GPTBotUser-agent: OAI-SearchBotUser-agent: ClaudeBotUser-agent: Google-ExtendedUser-agent: PerplexityBotAllow: /Read on 6 October 2026. The wildcard group disallows 24 paths, including /api/, /app/ and /sso/. The nine AI crawlers in the second group get none of them: their group says Allow: /, and that is all they read. Whether that is what was intended is not for us to say — it may well be — but a parser has to say it is what the file does.
2. Groups that name the same crawler are one group
RFC 9309 §2.2.1: if several groups name the same crawler, their rules are combined. This is the rule our validator got wrong. It took the first group that matched, which is what most parsers written from memory do, because the RFC's predecessor never said otherwise.
User-agent: GPTBotDisallow: /private/User-agent: GPTBotAllow: /private/help/GPTBot's rules here are both lines. /private/data is blocked and /private/help/faq is open. A first-group parser says /private/help/ is blocked; a last-group parser says /private/data is open. Both are wrong, in opposite directions, on a file that looks like it was edited twice by two people — which is most files older than a year.
3. The longest matching path wins; a tie goes to Allow
Within the rules a crawler reads, order does not matter. The rule whose path matches the most characters of the URL wins, and when an Allow and a Disallow match the same length, Allow wins. A Disallow: / followed by Allow: / is therefore an open site, and a file that says Allow: / then Disallow: /* is a closed one, because /* is two characters.
User-agent: *Disallow: /blogAllow: /blog/public/blog/public/a is open and /blog/private is blocked, and swapping the two lines changes nothing.
4. A group ends after a rule, not at a blank line
A group is one or more User-agent lines followed by rules. It ends when, after at least one Allow or Disallow, the next User-agent line begins. Blank lines do not end it. Comments do not. Neither does Sitemap:, nor Crawl-delay:, nor any line the parser does not recognise. The consequence is a trap we had not seen described anywhere before the parity run turned it up.
User-agent: baiducrawl-delay: 1User-agent: *Allow: /*?tab=achievements&achievement=*Disallow: /*/*/pulseRead on 6 October 2026. The baidu group contains only a Crawl-delay, which is not a rule, so the group has not ended when User-agent: * arrives two blank lines later. Google reads one group naming both, and baidu gets every wildcard rule. Here that is harmless, and probably what GitHub wanted. Put a User-agent: Googlebot with only a Crawl-delay above an AI section that says Disallow: /, and Googlebot is locked out of the site. Python's urllib.robotparser ends groups at blank lines and reads the same file the other way.
Three parser mistakes that look fine on most files
Matching the crawler name as a substring
A User-agent value is a whole token, compared case-insensitively. Applebot does not name Applebot-Extended; Claude does not name ClaudeBot. Our old parser matched substrings, and it cost us two wrong verdicts on live sites: amazon.com has a group for omgili, which we read as a block on Omgilibot, and en.wikipedia.org has one for Fetch, which we read as a block on Meta's meta-externalfetcher. Both crawlers are in fact unnamed on those sites and read the wildcard group. (Both groups are still there on 6 October 2026.)
Taking the first group that matches
The combined-groups rule above, seen in the wild. This one costs a verdict in the dangerous direction.
User-Agent: *User-Agent: ChatGPT-UserUser-Agent: OAI-SearchBotUser-Agent: Google-ExtendedUser-Agent: PerplexityBotDisallow: /AddListingDisallow: /ShowUserReviews# … 650 more lines of this group …User-Agent: Google-ExtendedDisallow: /Allow: /Restaurants-g55711-Dallas_Texas.htmlRead on 6 October 2026. Google-Extended, the token that controls whether pages feed Gemini, is named near the top in a group it shares with * and seven other crawlers, and again at the bottom with Disallow: / and one Allow. A first-group parser reports the top group and calls Google-Extended free to read nearly the whole site. Google combines both and reads Disallow: /: the crawler is shut out of everything except the paths an Allow line names, one of which is a single Dallas restaurants page. Our checker gave the friendly answer for the better part of a month.
Splitting groups at blank lines
The github.com case above. It reads correctly on almost every file, because almost every group has a rule in it. Then a group with only a Crawl-delay or a Sitemap line arrives, and the verdict flips without anyone noticing.
Where Google reads more than the RFC, and the one place we do not follow
Google's parser is forgiving about how a line is written. A byte-order mark before the first line is skipped. A bare carriage return counts as a newline. Dissallow:, Useragent: and a few other misspellings it has seen often enough are read as the directive they meant. A missing colon is read if there is whitespace where it should be. We read every one of these the same way, and flag the line, because a crawler that is stricter than Google will read the file differently and the author should know.
One Google-only rule we deliberately do not copy: Allow: /index.html is read by Google as also allowing /, so that a homepage is not blocked by its own Disallow: /. The RFC does not say this, and nothing says the AI crawlers do it. Copying it would mean telling a site its homepage is open to crawlers that may well be reading Disallow: / there. On every rule where we have a choice, we choose the reading that errs towards “blocked”, because “you are fine” is the one wrong answer that guarantees the rest of the work is wasted.
Thirty large sites, five wrong verdicts
After the fix we re-read the robots.txt of 30 large sites with both parsers. The old one had been wrong on five of them: tripadvisor.com (first-group), amazon.com and en.wikipedia.org (substring), and two where a Crawl-delay-only group had joined the one below it — github.com and, again, en.wikipedia.org. On every one of the five the error was in the friendly direction. That is the expected shape: a careless parser is not wrong at random, it is wrong towards “allowed”, because the shortcuts it takes all drop rules.
The robots.txt validator now prints, beside every line, the group it landed in and the crawlers that group covers, and when a named group releases a crawler from the wildcard rules it lists which paths just opened. The AI crawler checker applies the same parser to a live file. If either disagrees with Google's binary on a file of yours, send us the file: the fixture gets a case and the test gets a verdict that is Google's to give, not ours.
Questions
- Do the AI crawlers parse robots.txt the way Google does?
- Nobody outside those companies can check. What we can say: Google's library was written by the authors of RFC 9309, the standard every AI operator says it follows, and on the cases where Google reads more than the RFC strictly asks, we read the file the same way and flag the line. The one Google-only rule we did not copy errs on the friendly side, which is the dangerous side for this job.
- Does the order of groups matter?
- Not for which group applies: a crawler collects every group that names it, wherever they sit, and reads them as one. Not for which rule wins either: the longest matching path wins regardless of position. A parser that takes the first group it finds, or the first rule that matches, will be right on most files and wrong on exactly the ones that were edited twice.
- Does a blank line end a group?
- No. A group ends when a rule line (Allow or Disallow) has been read and the next User-agent line begins. Blank lines, comments, Sitemap lines and Crawl-delay do not end one. Python's urllib.robotparser ends a group at a blank line, which is why it reads the Crawl-delay case on this page the opposite way from Google.
- What does Crawl-delay do?
- Google ignores it. A few crawlers honour it — Anthropic documents it for ClaudeBot — and where they do it only slows them down; it never keeps one out. Its real effect on a file is the one above: it does not end a group, so a group with nothing but a Crawl-delay in it swallows the rules of the next one.
- My file names a bot and * in the same group. Which rules does the bot get?
- That group's, plus any other group that names it. That is the tripadvisor shape below: Google-Extended is in a nine-agent group at the top and in its own group at the bottom, and it gets both. What it does not get is a group that names only *, because once anything names a crawler, the * group is no longer about it.
- How do I check my own file?
- Paste it into the robots.txt validator. It reads the file the way described here, prints the group each line landed in, flags the lines that do nothing, and tests a path against Googlebot, Bingbot and every AI crawler. If it disagrees with Google's parser on a file, that is a bug and we would like the file.
See who the engines recommend instead
12 buyer questions across ChatGPT, Claude and Gemini. You get the questions your client loses, who won them, and what is missing from the site. No score.