
# Which AI Fetchers Send Which Headers, Measured on a Live Site

<p class="post-orient">A header-by-header record of what ten AI fetchers sent to one Cloudflare Worker on September 3, 2026, checked against each vendor's own documentation, IP lists and signing keys, with the raw captures published.</p>

When a person asks an AI assistant to read a web page, the assistant sends a request to that page's server. Whether the site owner can tell that request apart from a human visitor depends entirely on what the assistant chooses to put in it. I asked ChatGPT, Claude, Gemini, Grok, Perplexity, Copilot, Mistral, DuckDuckGo, Claude Code and OpenAI's Codex to each open a unique URL on this site and recorded every header the server saw. Five identify themselves in a way you can verify against a vendor list. One, DuckDuckGo's, goes further and signs every request with a published key, the only fetcher in the set that proves who it is. One identifies itself with a single word that appears in no documentation. Three reported a result about the page without ever requesting it. And one does not identify itself at all: Grok fetched the page eight times in twelve seconds from eight networks on four continents, including a mobile carrier in Ireland and a home ISP in Brazil, wearing Safari and Chrome headers that pass every browser check my analytics have. A second run an hour later produced eight more, from eight different networks. If you count readers on your own server, those sixteen requests were sixteen people.

The findings, each explained below:

- **ChatGPT-User** is the best-behaved classic fetcher: documented User-Agent, published IP range that the request matched, browser-shaped `Accept` and `Accept-Language`, over HTTP/2 from Microsoft's network.
- **Claude-User** identifies itself and its IP matches Anthropic's published list, but the request carries almost nothing else: `Accept: */*`, no language, HTTP/1.1, from Google Cloud.
- **MistralAI-User** matches its documentation and IP list and sends the most complete browser header set of any declared fetcher, Fetch Metadata included, from Azure.
- **DuckAssistBot** is the only fetcher that signed: a Web Bot Auth signature whose key id matches the Ed25519 key at its well-known directory. Its documentation does not mention this.
- **Gemini** sends `User-Agent: Google` and `Accept: */*` from a Google address that is in none of Google's five published crawler and fetcher IP lists, and "Google" matches none of the twelve user-triggered fetchers Google documents.
- **Grok** on the web sends no token, no signature, and full browser headers from rotating proxy exits, several of them residential or mobile. There is no request fact that separates it from a person. The pattern reproduced exactly on a second run. Logged in, on the phone app and on the website alike, Grok used a different path entirely: a self-declared `HeadlessChrome/148` on Google Cloud that rendered the page and fetched its stylesheet and logo.
- **Perplexity** reported "HTTP 200 OK" and the correct heading for a real page without any request reaching the origin, and reported a fetch failure for an unknown path, also without a request. Logged in, it did the same, and instead dispatched its search crawler to `robots.txt`, `/about` and `/essays`.
- **Copilot**, logged in, reported twice that its fetch tool "returned an empty result". No request reached the origin, and Cloudflare's edge firewall log shows nothing was blocked.
- **Claude Code's fetch tool** runs on the user's own machine, not on Anthropic's, and asks for Markdown before HTML.
- **Codex's web search** answered correctly about the page without requesting it; 76 minutes later an unnamed client on Amazon fetched that exact URL.

Three of these matter beyond this site. The Grok result means that server-side "human versus bot" counts on any site are inflated by an unknowable amount whenever people use Grok to read pages. The Gemini result means the largest search company in the world runs a consumer fetcher that its own crawler documentation does not describe. The DuckDuckGo result means the mechanism that would fix both already runs in production at a mainstream assistant, and the others have simply not adopted it.

## The captures

Each assistant was given a fresh chat and asked to open a URL unique to it, either `https://gkoreli.com/does-llms-txt-work?probe=<name>` or, in the second run, `https://gkoreli.com/probe/<name>`, and quote the heading. A request carrying that URL can only have come from that assistant. The second form returns a 404, which the Worker logs like any other request; it was used to test whether fetchers that failed on the query-string form would fetch a plain path. The Worker in front of the site was tailed with `wrangler tail --format json`, which records every request header plus the network facts Cloudflare attaches: autonomous system, country, HTTP version, and TLS ClientHello fingerprints. The full captures are in the [research directory](https://github.com/gkoreli/blog/tree/main/packages/blog/drafts/research/ai-fetcher-headers), and the requests are also published as data: [`captures.jsonl`](https://github.com/gkoreli/blog/blob/main/packages/blog/drafts/research/ai-fetcher-headers/data/captures.jsonl) and [`captures.csv`](https://github.com/gkoreli/blog/blob/main/packages/blog/drafts/research/ai-fetcher-headers/data/captures.csv), one record per request with every header as received, the network facts, and the vendor-list match, alongside the vendor IP lists and the Cloudflare firewall events as fetched that day. Client addresses appear only where they fall inside a vendor's published list. This section is the summary.

| Fetcher | User-Agent (as received) | `Accept` | `Accept-Language` | Fetch Metadata | Network | IP in vendor list | HTTP |
|---|---|---|---|---|---|---|---|
| ChatGPT | `Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; ChatGPT-User/1.0; +https://openai.com/bot` | browser-shaped, 8 types | `en-US,en;q=0.9` | none | AS8075 Microsoft | yes, `chatgpt-user.json` | HTTP/2 |
| Claude.ai | `Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; Claude-User/1.0; +claude-user@anthropic.com)` | `*/*` | absent | none | AS396982 Google Cloud | yes, `claude.com/crawling/bots.json` | HTTP/1.1 |
| Mistral (Vibe), 2 requests | `Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; MistralAI-User/1.0; +https://docs.mistral.ai/robots)` | browser-shaped, plus `application/json` | `en-US,en;q=0.9` (first request) | `navigate` / `document` / `none`, `Sec-Fetch-User: ?1`, `Sec-CH-UA`, `Cache-Control: no-cache` (first request); none (second) | AS8075 Microsoft | yes, `mistralai-user-ips.json` | HTTP/2 |
| DuckDuckGo duck.ai | `DuckAssistBot/1.2; (+http://duckduckgo.com/duckassistbot.html)` | `*/*` | absent | none, but `Signature`, `Signature-Input`, `Signature-Agent: "https://assistbot.duckduckgo.com"` | AS8075 Microsoft | yes, `duckassistbot.json`; signature key verified | HTTP/2 |
| Gemini | `Google` | `*/*` | absent | none | AS15169 Google | no (checked 5 lists) | HTTP/1.1 |
| Grok, run 1, 8 requests | Safari 26.2 on macOS (2), Chrome 143 on macOS (5), Chrome 142 (1) | browser-shaped | `en-US,en;q=0.9` | `navigate` / `document` / `none` | 8 ASNs: 5 hosting or transit, 1 mobile carrier, 2 consumer ISPs (registry names) | no | HTTP/2 |
| Grok, run 2, 8 requests | Safari 26.2 (4), Chrome 143 (2), Chrome 142 (2) | browser-shaped | `en-US,en;q=0.9` | `navigate` / `document` / `none` | 8 different ASNs: 5 hosting or transit, 3 telecoms (two Brazilian, one Argentine) | no | HTTP/2 |
| Grok, logged in (phone app and website), 1 to 2 renders + subresources, plus one bare `Mozilla/5.0` client | `Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) HeadlessChrome/148.0.0.0 Safari/537.36` | browser-shaped | `en-US,en;q=0.9` | `navigate` / `document` / `none`, `Sec-Fetch-User: ?1`, `Sec-CH-UA` platform Linux | AS396982 Google Cloud | no list exists | HTTP/2 |
| Perplexity, 7 attempts | no request for any probe URL; PerplexityBot crawled 3 other paths | `PerplexityBot/1.0` with `From: crawler-support@perplexity.ai`, no `Accept` | absent | none | AS14618 Amazon | yes, `perplexitybot.json` | HTTP/1.1 |
| Copilot, 2 attempts | no request reached the origin | no request | no request | no request | no request | no request | no request |
| Claude Code | `Claude-User (claude-code/2.1.259; +https://support.anthropic.com/)` | `text/markdown, text/html, */*` | absent | none | my own ISP | not applicable, runs locally | HTTP/1.1 |
| Codex CLI web search | none at answer time; 76 min later `Mozilla/5.0 (compatible)` fetched the exact URL | `*/*` | absent | none | AS14618 Amazon | no | HTTP/1.1 |
| Chrome, human baseline | `Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/152.0.0.0 Safari/537.36` | browser-shaped | `en-US,en;q=0.9,ka;q=0.8,ru;q=0.7` | `navigate` / `document` / `cross-site`, plus `Sec-CH-UA`, `Upgrade-Insecure-Requests` | my own ISP | not applicable, a person | HTTP/3 |

"Fetch Metadata" means the `Sec-Fetch-Mode`, `Sec-Fetch-Dest` and `Sec-Fetch-Site` headers that every current browser engine sends on a navigation. Of the declared fetchers only Mistral's first request sends them, which suggests a real browser engine behind that fetch. Grok's requests all do.

Every assistant in the table was reached. Copilot demands a sign-in and was probed from my own account; ChatGPT, Claude and Gemini also used logged-in sessions; the rest were used anonymously. For every "no request" row I also pulled Cloudflare's firewall event log for the zone over the same hours to make sure the edge had not blocked the fetch before it reached the Worker. It had not: the only blocks in that window were exploit scans against the homepage.

## ChatGPT-User: the reference behaviour

OpenAI's fetcher is what every other fetcher should be measured against. The User-Agent matches OpenAI's [published string](https://developers.openai.com/api/docs/bots) character for character. The source address, `9.129.45.186`, is inside the range OpenAI publishes at `openai.com/chatgpt-user.json`. The `Accept` header is the same list Chrome sends, `Accept-Language` is present, and the request came over HTTP/2 with an Envoy timeout header of 15 seconds, which tells you roughly how long OpenAI is willing to wait for your page.

The ChatGPT mobile app uses the same fetcher: a probe sent from the iPhone app produced one request with an identical User-Agent and header set, again from an address in OpenAI's list. That probe pointed at a 404 page, and ChatGPT reported the status correctly but said no body or heading was returned. The server sent a full 404 page with a heading; the tool discarded it. DuckDuckGo's tool did the same, while Grok and Mistral read the 404 page's heading. How much of a non-200 response an assistant lets its model see varies by vendor. After the 404, the mobile app also tried to reach the URL with Python's `requests` library from its code sandbox, with the User-Agent `Mozilla/5.0`; that request never arrived, so the sandbox has no route to the open web. Two minutes later ChatGPT-User fetched four more posts on the site, apparently looking for the page. One question, six page loads. Asking the same question again a few minutes later produced the same 404 answer and no new request: ChatGPT reused the earlier result rather than fetching twice.

Two things to know. First, OpenAI states plainly that "because these actions are initiated by a user, robots.txt rules may not apply" to ChatGPT-User. Second, after fetching the probe URL, ChatGPT also fetched the homepage three seconds later, unprompted. One question produced two page loads.

For a site owner, ChatGPT-User is fully nameable: token plus IP range is a fact, not a guess, and my analytics label it as such.

## Claude-User: named, verifiable, and nearly silent

Anthropic's fetcher from claude.ai also identifies itself and also came from an address in the vendor's published list. Everything else about the request is minimal: `Accept: */*`, no `Accept-Language`, HTTP/1.1, from a Google Cloud address. Anthropic's [documentation](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler) names the `Claude-User` token and its purpose but does not print the full User-Agent string; I could not find the `Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; ...)` form above anywhere in Anthropic's documentation, so the table is the record.

One difference from OpenAI is worth noting. Anthropic says all its bots, Claude-User included, "respect 'do not crawl' signals by honoring industry standard directives in robots.txt". OpenAI, Perplexity and Google all say the opposite for their user-triggered fetchers: robots.txt "may not apply" or is "generally ignored" because a person asked. Anthropic is the only one of the four that documents robots.txt as binding on its on-demand fetcher.

Claude Code's fetch tool is a different animal that happens to share the token. Its request came from my own residential IP, because the tool runs inside the CLI on the user's machine. It sent `Accept: text/markdown, text/html, */*`. That is the first request I have seen in this site's logs that asks for Markdown before HTML. At capture time this site's Markdown content negotiation was merged but not yet deployed, so it got HTML. The deploy went out later the same day; the same request now gets the post's Markdown source, and the observation row says so.

## MistralAI-User: a browser behind the token

Mistral's assistant, now branded Vibe, produced two requests for its probe URL within the same second, both from Azure addresses inside Mistral's [published list](https://docs.mistral.ai/robots) and both carrying the documented `MistralAI-User/1.0` User-Agent. The first looks like a real browser engine: Chrome's `Accept` list with `application/json` added, `Accept-Language`, the full Fetch Metadata set including `Sec-Fetch-User: ?1`, `Sec-CH-UA` client hints, and `Cache-Control: no-cache`. The second is a bare HTTP client with an `Accept` list and, oddly, a `Content-Length` header on a GET. Two fetch paths, one token, one question.

Mistral's documentation says the fetcher "handles user actions in Vibe" and is "not used for crawling the web in any automatic fashion, nor to crawl content for generative AI training". It does not say whether robots.txt binds it. For a site owner it is fully nameable: token plus IP list.

## DuckAssistBot: the only fetcher that proves who it is

DuckDuckGo's duck.ai fetched its probe URL once, from an Azure address inside DuckDuckGo's [published list](https://duckduckgo.com/duckassistbot.html), with the documented `DuckAssistBot/1.2` User-Agent and `Accept: */*`. Then it did something no other fetcher in this study did. It signed the request:

```text
signature-agent: "https://assistbot.duckduckgo.com"
signature-input: sig1=("@authority" "signature-agent");created=1788411572;expires=1788412172;keyid="Ov3HDsa8JQ39dPEYFvFFN-cUpnz9yNI8LDvr-5LeiBM";alg="ed25519";tag="web-bot-auth"
signature: sig1=:NBrgbVdKFeZEFLnpQx0osM5xAZ5wfGP1TBvYC2NBrYycNLKuX7EU+lsLxylIYn8A0f3zshL8IZRdP+fj3VkzBA==:
```

This is [Web Bot Auth](https://datatracker.ietf.org/doc/draft-ietf-webbotauth-httpsig-protocol/), the IETF draft that applies RFC 9421 HTTP Message Signatures to automated traffic. The `Signature-Agent` header names a host; that host serves a key directory at `/.well-known/http-message-signatures-directory`; the `keyid` is the RFC 7638 thumbprint of the signing key. I fetched DuckDuckGo's directory. It holds one Ed25519 key, its thumbprint is `Ov3HDsa8JQ39dPEYFvFFN-cUpnz9yNI8LDvr-5LeiBM`, and that is the key id in the captured request. The directory response is itself signed. A Worker with `crypto.subtle` can verify the signature over `@authority` and `Signature-Agent` and know, not infer, that the request came from DuckDuckGo. This site's Worker now does exactly that: a second DuckAssistBot fetch of this article, after the deploy, was stored as a signed agent with the signature verified against the fetched key, the first request on this site to carry a proven identity. Cloudflare Radar lists DuckAssistBot in its [bot directory](https://radar.cloudflare.com/bots/directory/duckassistbot).

DuckDuckGo's search crawler, DuckDuckBot, signs as well, but writes the `Signature-Agent` value as a bare unquoted URI, a form the current draft has deprecated, and my Worker's parser rejected it; the fix is queued. DuckDuckGo's own documentation for the bot does not mention signing at all. The one vendor doing the right thing is not advertising it. Among AI assistants, the public record of who signs named OpenAI's ChatGPT agent and Google's Google-Agent; DuckAssistBot belongs on that list.

## Perplexity: answers without requests

Perplexity was given seven chances, five anonymous and two from a logged-in account. Anonymously: three times it was asked to open a query-string probe URL and reported that "the page could not be retrieved"; once it was asked to open a plain 404 path and reported "Failed to fetch content"; once it was asked to open a real, existing post with no query string, and it answered "The page returned successfully (HTTP 200 OK)" with the correct H1. The origin saw none of the five. The "200 OK" came from Perplexity's index, and the tool presented an index hit as a live fetch with a status code.

Logged in, the two probes went the same way: the unknown path "still cannot be retrieved" and the real post answered correctly, with neither URL requested. But this time something did arrive. Within the same second, three requests from Amazon addresses inside Perplexity's [published `PerplexityBot` range](https://docs.perplexity.ai/guides/bots) fetched `/robots.txt`, `/about` and `/essays`, each carrying the `PerplexityBot/1.0` User-Agent and a `From: crawler-support@perplexity.ai` header. Perplexity's reply said it had "searched for the exact path". So a logged-in user asking for one URL triggered the search crawler against three other pages of the site, robots.txt first, and never the page that was asked for. Attribution here is by timing: PerplexityBot also crawls this site on its own schedule, but a three-request burst starting with robots.txt in the second the prompt was sent is not its routine pattern.

`Perplexity-User`, the fetcher Perplexity documents for exactly this situation, did not appear in any of the seven attempts. It is nameable by token and IP list when it does appear; on this site, for these prompts, it did not.

## Copilot: an empty result, twice

Microsoft Copilot, signed in, was asked twice to open its probe URL. Both times it said its fetch tool "returned an empty result" with "no HTTP status code, no HTML, no text, no error message", and it offered the usual explanations: the server blocks automated requests, the page needs JavaScript, the response type could not be parsed. None of those happened. The origin received no request at all, and the zone's edge firewall log for that hour contains no block, challenge, or rate limit for any Copilot, Bing, or Microsoft address. Whatever Copilot's fetch tool does with a URL it has never indexed, it does not fetch it. Microsoft documents `bingbot` and `MicrosoftPreview` but, as far as I can find, no on-demand fetcher for Copilot chat, which is consistent with what the log shows.

## Gemini: one word, no documentation

The Gemini app fetched the page with `User-Agent: Google`. Not `Googlebot`, not `Google-Agent`, not `Google-GeminiNotebook`. Just the company name.

Google maintains a [page listing its user-triggered fetchers](https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers): twelve of them, each with a full User-Agent string and an IP range file. `Google` on its own is not among them. The source address, in Google's own AS15169, is not in `googlebot.json`, `special-crawlers.json`, `user-triggered-fetchers.json`, `user-triggered-fetchers-google.json` or `user-triggered-agents.json`. I checked all five on the day of the capture.

So the fetcher behind the consumer Gemini app is undocumented on the page Google wrote for exactly this purpose. A site owner sees a request from a Google address with a one-word User-Agent and has no vendor statement to match it against. My analytics can record the literal token `Google` as a fact. They cannot call it Gemini, because Google has not said so.

## Grok: indistinguishable from people, by design

This is the finding that changes what server-side analytics can claim.

Grok, used anonymously in "Fast" mode, showed a tool step reading "Opened page gkoreli.com/does-llms-txt-work?probe=grok" and quoted the H1 correctly. The origin saw eight requests for that URL between 04:03:18 and 04:03:30 UTC:

| Time (UTC) | ASN and registry holder | Cloudflare's organisation string | Country (Cloudflare) | Claimed browser |
|---|---|---|---|---|
| 04:03:18 | AS9009 M247 Europe SRL (hosting, Romania) | Aventice LLC | US | Safari 26.2, macOS |
| 04:03:18 | AS3257 GTT Communications (transit) | Web2Objects LLC | US | Safari 26.2, macOS |
| 04:03:18 | AS132817 DZCRD Networks Ltd (ISP, Bangladesh) | DZCRD Networks Ltd | Netherlands | Chrome 143, macOS |
| 04:03:18 | AS13280 Three Ireland (Hutchison), mobile subscriber pool | Three Ireland (Hutchison) - Mobile Subscriber Pools | Ireland | Chrome 143, macOS |
| 04:03:18 | AS262988 Pombonet Telecomunicações (ISP, Brazil) | Pombonet Telecomunicações e Informática | Brazil | Chrome 143, macOS |
| 04:03:19 | AS212238 Datacamp Limited (hosting and CDN, UK) | Private Customer | South Africa | Chrome 143, macOS |
| 04:03:22 | AS7979 Servers.com (hosting) | Servers.com, Inc. | US | Chrome 142, macOS |
| 04:03:30 | AS398781 Oculus Networks Inc (hosting) | Private Customer | US | Chrome 143, macOS |

Registry holders are from RIPEstat and Team Cymru, checked on the day of the capture. Cloudflare attaches its own organisation string to each request; for the proxy exits it names the company registered for the IP block rather than the network that announces it, so the two columns disagree wherever a hosting provider resells address space. Both are recorded in the research directory.

Every request carried a complete browser header set: a Safari or Chrome User-Agent, Chrome's exact `Accept` list, `Accept-Language: en-US,en;q=0.9`, `Sec-Fetch-Mode: navigate`, `Sec-Fetch-Dest: document`, `Sec-Fetch-Site: none`, `Priority: u=0, i`, and on the Chrome-claiming requests the `Sec-CH-UA` client hints with matching version numbers. The words "Grok" and "xAI" appear nowhere. There is no `Signature-Agent` header. The TLS fingerprints differ from request to request, so this is not one client behind eight addresses; it is several client stacks.

An hour later I asked Grok again, anonymously, for a different unique URL. Eight requests again, in six seconds, from eight networks: GTT and Datacamp appeared in both runs, and the other six were new: Claro and DESTAK NET in Brazil, ARSAT, the Argentine state telecom, and three hosting providers, B2 Net Solutions, Ace Data Centers, and code200 in Lithuania. Same Safari and Chrome header sets, same absence of any token.

A third probe, sent from the Grok iPhone app while signed in, took a different path altogether. One request arrived from a Google Cloud address with the User-Agent `HeadlessChrome/148.0.0.0` on Linux, `Sec-Fetch-User: ?1`, Chrome 148 client hints, and then, a second later, same-origin requests for `/main.css` and the site logo with the probe page as referer. That is a real Chromium rendering the page, and it says so in its User-Agent: Chromium inserts the `Headless` prefix whenever it runs without a display. A fourth probe from grok.com on the desktop, also signed in, did the same: two headless renders six seconds apart, each pulling the stylesheet and logo, followed by one bare request with the User-Agent `Mozilla/5.0` and nothing else, also from Google Cloud. The eight-exit proxy pattern did not appear in either signed-in run.

So Grok has at least two fetch paths, and the split is by account, not by device: anonymous sessions get an undeclared proxy pool wearing consumer browser headers, signed-in sessions get a declared headless Chromium on Google Cloud. The second is easy to label as automation on a cloud host. The first is not, and it is what anyone using Grok without an account gets.

Attribution rests on the probe string: nobody but Grok was ever given any of the three URLs, and the requests arrived within thirty seconds of each prompt. The pattern is not new. A [February 2026 experiment by Stackfox](https://stackfox.co/research/grok-user-agent) used the same method and saw Chrome and iPhone Safari User-Agents from two rotating proxy providers, M247 Europe and Datacamp. Both appear in my captures. What this capture adds is that the exits now include a mobile carrier's subscriber pool and consumer ISPs on four continents, and that a single question produces eight fetches, not one. xAI publishes no fetcher documentation, no User-Agent token, and no IP list that I or the other researchers who have looked could find.

For counting purposes: my Worker classified all eight as browsers, and it was right to. The rule it applies, described in [How I Built First-Party Analytics for a Personal Blog](/first-party-analytics-for-a-personal-blog) and refined since, is that a request with a browser User-Agent, navigation-shaped Fetch Metadata, and a network that is not a known hosting provider is recorded as a browser. Three of the eight requests in the first run, and three in the second, came from consumer networks that no hosting list would ever contain: a mobile carrier, two Brazilian ISPs, a Bangladeshi ISP, an Argentine telecom. The rule cannot separate them from readers, and neither can any other server-side rule I know of. Grok's traffic is a human-shaped hole in every server-side audience count on the web.

## Codex: the fetch that never happened

I ran OpenAI's Codex CLI with its web search tool enabled and asked it to open the probe URL. It printed a search step with the URL, encoded as `probe=codex%2Dsearch`, then answered with the exact H1 and a link to the page. At answer time the origin had received no request for that URL. Neither the tail nor the database had one.

The page's content came from OpenAI's search index, which had the article from earlier crawls, and the tool presented that as having opened the URL. For the reader the answer was correct. For the site owner it was a read that no server log showed. Then, 76 minutes later, one request for exactly that URL arrived, with the hyphen still percent-encoded the way Codex had printed it, from an Amazon address, with the User-Agent `Mozilla/5.0 (compatible)`, `Accept: */*`, and no token, contact or signature. That URL existed nowhere but Codex's search step. Whatever fetched it is part of OpenAI's search pipeline catching up on a URL it had answered about without visiting, and it identified itself as nothing. Reads that happen this way are invisible to first-party analytics when they happen, and unattributable when the fetch finally comes.

## What the open-source detectors make of these requests

Most sites do not write their own classifier. They use a library. So I ran the six captured User-Agent strings through the detectors the open-source world actually ships: [isbot](https://github.com/omrilotan/isbot) 5.2.2 (the npm package), Matomo's [device-detector](https://github.com/matomo-org/device-detector) bot list, [crawler-user-agents](https://github.com/monperrus/crawler-user-agents), the [ai.robots.txt](https://github.com/ai-robots-txt/ai.robots.txt) registry, and GoatCounter's [isbot](https://github.com/arp242/isbot) rules, all at their current `main` on the day of the capture.

| Captured User-Agent | isbot | device-detector | crawler-user-agents | ai.robots.txt |
|---|---|---|---|---|
| ChatGPT-User | bot | `ChatGPT-User` | `ChatGPT-User` | listed, OpenAI |
| Claude-User | bot | `Claude-User` | `Claude-User` | listed, Anthropic |
| `Google` | bot, generic `google` pattern | `Googlebot`, "Search bot" | no match | no entry |
| Claude Code | bot | `Claude-User` | `Claude-User` | `Claude-Code`, operator "unclear" |
| Grok, Safari UA | not a bot | no match | no match | no entry |
| Grok, Chrome UA | not a bot | no match | no match | no entry |

Three things fall out. The declared fetchers are named identically everywhere; on ChatGPT-User and Claude-User the open-source consensus and my Worker agree. The bare `Google` string is known to exactly one detector, and that detector files it under Googlebot as search crawling, by a rule added in August 2023, months before the Gemini app existed. So every Matomo installation counts a person reading through Gemini as Google indexing the page. And Grok passes every check, because there is nothing to match: ai.robots.txt, whose whole purpose is to let sites express AI-agent policy in robots.txt, has no xAI entry and cannot have one. GoatCounter's IP-range rules would catch one of the eight exits, the one at Servers.com. Isbot's own pattern for Claude Code, `^claude-code/`, does not match the real string either, which starts with `Claude-User`; the request is caught by a different pattern. Each of those is a pull request I will open with this capture as the evidence, and the results table lives in the research directory with the exact files checked.

The captured strings were also compared with the community's own records. crawler-user-agents stores the exact User-Agents its contributors have seen: ChatGPT-User, MistralAI-User, DuckAssistBot and PerplexityBot match those records byte for byte, which corroborates both sides. The recorded Claude-User instance differs from mine only in the case of the contact token. `Google` has no record anywhere. And the Grok exits were checked against X4BNet's open datacenter list of about forty-three thousand ranges: it flags three of the eight first-run exits and two of the eight second-run exits, none of them the consumer networks. Finally, the well-known signing-key directory was requested on seventeen vendor hosts; only chatgpt.com and assistbot.duckduckgo.com serve one.

## What this means if you run a site

This section is for anyone who counts visitors on their own server, whether with a Worker, a log parser, or a hosted product that runs at the edge.

**Four of the fetchers can be named as facts.** ChatGPT-User, Claude-User, MistralAI-User and DuckAssistBot publish both a token and an IP list, and every captured request matched its list. Token plus IP match is a verifiable identity. Label it, count it, and do not block it: these requests are a person reading your page through a tool. Perplexity-User is nameable the same way when it appears.

**One can be proven.** DuckAssistBot's signature verifies against a published key. That is a stronger fact than any IP list, because it survives a change of hosting and cannot be produced by a proxy pool. If you run an edge function, verifying it costs one cached key fetch and one Ed25519 check.

**One can be recorded but not named.** `User-Agent: Google` from a Google address is a fact. "Gemini" is an inference until Google documents it. Store the token as it arrived.

**One cannot be found.** Grok's requests will sit inside your browser count, and inside Cloudflare's, Plausible's, GoatCounter's and everyone else's, until xAI either adds a token or signs its requests. Web Bot Auth is the mechanism that would settle this, and DuckDuckGo's capture shows it is not theoretical: OpenAI's ChatGPT agent, Google's Google-Agent and DuckAssistBot sign today. Anthropic, Perplexity, Mistral and xAI do not. A verified signature is the only header fact that cannot be spoofed by a proxy pool; everything else in this article can be.

**Some reads leave no trace at all.** Index answers like Codex's and Perplexity's never reach you, even when the assistant reports a status code, and Copilot's tool reported a result for a page it never asked for. Your logs are a floor on AI readership, not a measurement of it.

**Do not trust your own labels until you have seen the raw requests.** The live version of my Worker at the time of the capture did not yet recognise the `Claude-User` token, so the Claude fetches were stored as browsers with `Accept: */*` and no language header, and it filed DuckAssistBot under `DuckDuckBot`, the search crawler. Both fixes were merged and deployed later the same day. Header-level captures like these are how you find out your classifier is wrong; aggregate dashboards never tell you.

## Method and limits

Captures were taken on 2026-09-03 in four runs, 03:55 to 04:05, 04:55 to 05:05, 05:13 to 05:15 and 05:20 to 05:24 UTC, with `wrangler tail --format json` against the production Worker for gkoreli.com. Each assistant received a prompt of the form "Please open this exact URL and tell me the exact text of its main heading: <URL>. Do not answer from memory; fetch the page." in a new chat. ChatGPT (Pro, web and iPhone app), Claude.ai (Max), Gemini, Copilot and the second Perplexity pair were used with logged-in accounts; Grok, Perplexity, Mistral and duck.ai were used anonymously. For every row with no origin request, the zone's firewall events (Cloudflare GraphQL `firewallEventsAdaptive`) for the surrounding hours were checked for blocks or challenges; none matched. Vendor documentation and IP lists were fetched the same day; the addresses were checked against them, and the captured strings were run through the open-source detectors, with TypeScript scripts that are in the research directory. The DuckDuckGo key thumbprint was recomputed from the published JWK.

What this does not show:

- **One or two requests each.** Most fetchers were probed once. Headers can vary by region, plan, model, or the tool the assistant chooses. Grok's twenty page requests came from four prompts.
- **Perplexity-User is absent from the origin.** Seven attempts, anonymous and logged in, produced no request for the asked URL. A page not yet in Perplexity's index, or a different prompt shape, might.
- **Copilot was probed twice from one account.** A different Copilot surface (Edge sidebar, Windows, Microsoft 365) may use a different tool.
- **Network names are registry holders.** Every ASN was checked against RIPEstat and Team Cymru on the day. Cloudflare's own `asOrganization` string differed for seven of the fourteen Grok exits; it names the IP block's registered organisation, not the announcing network. The first draft of this article used Cloudflare's strings and a reader caught the discrepancy.
- **Header order is lost.** The tail event delivers headers as a map. Order is a fingerprinting signal in its own right and is not analysed here.
- **Attribution for Grok is by timing and the unique URL**, not by any declaration from xAI. I consider eight requests for a URL that existed nowhere else, within thirty seconds of the prompt, twice, conclusive; a reader who wants stronger evidence can repeat the probe with their own URL.
- **The first DuckDuckGo signature was matched by key id only.** The second, after the deploy, was verified in full by the Worker (Ed25519 over `@authority` and `Signature-Agent`, key fetched from the directory, validity window checked) and stored as `signed-agent`, `verified`.

Evidence that would change the conclusions: a Google page documenting the `Google` User-Agent and its IP range; a Perplexity capture; any xAI documentation or a `Signature-Agent` header on a Grok request; a second round of probes showing different headers from the same vendors.

## What I am doing with this

The site's analytics will name what is declared and verified, record what is declared and unverified as the literal token, and stop pretending the remainder is anything more specific than "a browser-shaped request". The eight Grok fetches stay in the browser count with no asterisk, because there is no honest fact that would move them. That is uncomfortable for a site whose whole analytics project is about believable numbers, and the discomfort is the point: the number is a floor, and the article that reports it should say so.

The captures will be repeated monthly and this page updated in place with a dated changelog, so the table above stays a reference rather than a snapshot. If a vendor changes its headers, the change will show here first.

## Evidence ledger

**Captures, vendor documentation, IP lists, registries and open-source detectors checked:** September 3, 2026. Fetcher behaviour is volatile; the method, the raw captures and the data files are the durable layer.

| Claim | Source | Evidence date |
|---|---|---|
| ChatGPT-User User-Agent, IP list, "robots.txt rules may not apply" | [OpenAI bots documentation](https://developers.openai.com/api/docs/bots); `openai.com/chatgpt-user.json` | 2026-09-03 |
| Claude-User purpose, robots.txt honoured, IP list | [Anthropic crawler article](https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler); `claude.com/crawling/bots.json` | 2026-09-03 |
| Perplexity-User and PerplexityBot strings, "generally ignores robots.txt", IP lists | [Perplexity bots guide](https://docs.perplexity.ai/guides/bots); `perplexity.com/perplexity-user.json`, `perplexitybot.json` | 2026-09-03 |
| MistralAI-User string, purpose, IP list | [Mistral robots page](https://docs.mistral.ai/robots); `mistral.ai/mistralai-user-ips.json` | 2026-09-03 |
| DuckAssistBot string, purpose, IP list, no mention of signing | [DuckAssistBot page](https://duckduckgo.com/duckassistbot.html); `duckduckgo.com/duckassistbot.json` | 2026-09-03 |
| DuckAssistBot signing key | `https://assistbot.duckduckgo.com/.well-known/http-message-signatures-directory`, one Ed25519 key; RFC 7638 thumbprint recomputed | 2026-09-03 |
| Web Bot Auth wire format and directory lookup | [draft-ietf-webbotauth-httpsig-protocol](https://datatracker.ietf.org/doc/draft-ietf-webbotauth-httpsig-protocol/); [cloudflare/web-bot-auth](https://github.com/cloudflare/web-bot-auth) | 2026-09-03 |
| Google's twelve user-triggered fetchers, "generally ignore robots.txt", no bare `Google` UA | [Google user-triggered fetchers](https://developers.google.com/search/docs/crawling-indexing/google-user-triggered-fetchers); [Google common crawlers](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers); five Google IP range files | 2026-09-03 |
| Earlier undocumented Gemini UA (`GeminiiOS`) | [ppc.land report](https://ppc.land/gemini-ios-app-traffic-revealed-through-undocumented-user-agent/), [Search Engine Roundtable](https://www.seroundtable.com/geminiios-google-user-agent-40333.html) | 2025-10-27, checked 2026-09-03 |
| Grok proxy behaviour, prior observation of M247 and Datacamp | [Stackfox, Grok user agent research](https://stackfox.co/research/grok-user-agent) | February 2026, checked 2026-09-03 |
| No xAI fetcher documentation | Searched x.ai, grok.com, xAI docs; third-party directories only | 2026-09-03 |
| Chromium adds the `Headless` prefix without a display | Chromium source, `components/embedder_support/user_agent_utils.cc` (`if (HasSwitch(kHeadless)) product.insert(0, "Headless")`) | checked 2026-09-02 in the readers-vs-bots research |
| ASN registry holders | [RIPEstat as-overview](https://stat.ripe.net/docs/02.data-api/as-overview.html), [Team Cymru IP-to-ASN](https://www.team-cymru.com/ip-asn-mapping); ARIN WHOIS for block organisations | 2026-09-03 |
| Cloudflare `request.cf` fields (`asn`, `asOrganization`, TLS digests) | [Cloudflare Workers request properties](https://developers.cloudflare.com/workers/runtime-apis/request/#incomingrequestcfproperties) | 2026-09-03 |
| No edge block of any probe | Cloudflare GraphQL `firewallEventsAdaptive`, zone gkoreli.com, 03:30 to 05:45 UTC; file `data/firewall-events-2026-09-03.json` | 2026-09-03 |
| Open-source detector verdicts | [isbot](https://github.com/omrilotan/isbot) 5.2.2, [device-detector](https://github.com/matomo-org/device-detector) `regexes/bots.yml` and node-device-detector 2.2.7, [crawler-user-agents](https://github.com/monperrus/crawler-user-agents), [ai.robots.txt](https://github.com/ai-robots-txt/ai.robots.txt), [GoatCounter isbot](https://github.com/arp242/isbot) | 2026-09-03 |
| device-detector's `^Google$` rule dates from 2023-08-07 | GitHub blame, commit a9f29e5 | 2026-09-03 |
| Datacenter IP list used for the Grok exits | [X4BNet lists_vpn, datacenter ipv4](https://github.com/X4BNet/lists_vpn), 42,797 ranges | 2026-09-03 |
| Raw captures, probe log, data files, scripts | [research directory](https://github.com/gkoreli/blog/tree/main/packages/blog/drafts/research/ai-fetcher-headers) | 2026-09-03 |

Previously in this series: [Does llms.txt Work?](/does-llms-txt-work) tested what agents are supposed to read, and [How I Built First-Party Analytics for a Personal Blog](/first-party-analytics-for-a-personal-blog) built the counter that this article is now checking.


## More in Measurement boundaries

1. [Does llms.txt Work? What a Live Implementation Revealed](/does-llms-txt-work.md)
2. [How I Built First-Party Analytics for a Personal Blog](/first-party-analytics-for-a-personal-blog.md)
3. Which AI Fetchers Send Which Headers, Measured on a Live Site (current)
4. [How I Separate Readers from Bots on a Static Blog Without JavaScript](/how-i-separate-readers-from-bots-without-javascript.md)


## Cite this

```bibtex
@misc{koreli2026which,
  author={Koreli, Goga},
  title={Which AI Fetchers Send Which Headers, Measured on a Live Site},
  year={2026},
  month={September},
  howpublished={gkoreli.com},
  url={https://gkoreli.com/which-ai-fetchers-send-which-headers},
  note={CC BY-NC-ND 4.0}
}
```

Koreli, Goga. (September 2026). "Which AI Fetchers Send Which Headers, Measured on a Live Site". gkoreli.com. https://gkoreli.com/which-ai-fetchers-send-which-headers
