Bottom line
- A page can be fetched by AI agents four hundred times and never cited. Agent traffic measures access, and access is only the first of three signals.
Agent traffic in a server log measures access. It says an automated client asked for a page. It does not say the page was used in an answer, and it does not say a human followed a citation to the site. Reported as one number, the three signals hide the finding, which is usually the gap between them.
What is agent traffic, and what is it measured against?
Two published platforms name it as a feature. Scrunch lists AI agent and bot traffic as included on its Core plan, and its product menu names a separate agent traffic view. Profound's homepage describes agent analytics that track AI-sourced traffic and attribution across domains. Read on 2026-09-28, both make agent traffic a headline capability.
It is measured against an outcome, and the outcome is usually assumed: that being fetched leads to being cited, and being cited leads to visits. Each arrow is a separate claim, and each has its own evidence.
What does a log actually record?
A request: a timestamp, a path, a status code and a user-agent string. At the time of writing, the major AI providers publish separate user-agent strings for training crawlers, search indexing, and fetches triggered by a person asking a question. Verify the current names in each provider's documentation, since they change. Those three classes mean different things.
A training crawl says the page may inform a model. An index crawl says it may appear in a retrieval set. A user-triggered fetch says someone asked a question that led the system to open the page now. A blended agent traffic count adds them together, and the sum has no single meaning. A user-agent string is also a claim the client makes about itself. For anything that will drive a decision, check the source address against the provider's published ranges, or the count includes impostors. Fetch counts inflate under retries and pagination as well: one question can trigger several fetches of the same page, so count distinct pages per day, not requests.
Illustrative lines:
2026-09-27T09:14:02Z GET /guides/pricing 200 UA=search-indexer
2026-09-27T09:14:40Z GET /guides/pricing 200 UA=user-fetcher
2026-09-27T09:15:11Z GET /guides/pricing 403 UA=training-crawler
The same page, three classes, three outcomes. Counting requests hides all of it.
What is the gap between access and citation?
Take a constructed corpus of twenty pieces over thirty days. Fourteen were fetched by an index crawler at least once. Nine were fetched by a user-triggered client at least once. Against a fixed set of sixty queries, six were cited. Every number is illustrative.
So fourteen were reachable, nine were opened during a live question, and six were used. The ratio of fetched to cited is more than two to one, and it is the ratio that carries information. The pieces that were fetched and never cited are the control: they were reachable and unused, which points to structure or relevance, not access.
Separate the effect of structure from the effect of subject. Among the fetched-but-uncited pieces, the common features are a summary buried below the fold and headings that are not questions. Among the cited, they are the opposite. Structure explains more of this than subject, which is an uncomfortable finding for an editorial team.
What does agent traffic miss?
Cached retrieval. An answer can cite a page from an index built weeks earlier without a fresh fetch, so a quiet week in the log does not mean a quiet week in answers. It also misses the human half: a person who reads a citation and clicks arrives as ordinary referral traffic, in a different table.
So the three signals live in three places. Agent traffic sits in the edge and server logs. Citations sit in the answer sampling, against a query set and a date. Human visits sit in analytics, and can land in direct traffic when a surface does not pass a referrer, so treat that column as a floor. A dashboard that shows one of them is showing a third of the picture, and one that adds them together is showing none.
What change is worth making, and what does it cost?
Report the three side by side for the same thirty-day window. The cost is small: a log query and a saved report, roughly two hours to build, and ten minutes a week to refresh. The value is the ratio between columns.
A quarterly review of the ratios, not the raw counts, is enough. Then attack the widest gap. If fetched is high and cited is low, restructure: a summary sentence, question-shaped headings, one claim per paragraph. If fetched is low, look at access first: the robots file, the edge rules, the render path.
Checkpoint GTM works on all three layers as one programme, so the team reading the logs is the team restructuring the pages and checking access, and the gap between columns has an owner. For a buyer who has the tile but not the hours to act on the ratio, that is the stronger arrangement. A team with an analyst and spare editorial capacity can run the same comparison from a platform's agent traffic view alone.
What measurement would confirm it worked?
Re-run the same sixty queries four weeks after restructuring, and re-read the log for the same twenty pieces. If cited rises among the restructured pieces and not among the untouched ones, structure was the cause. A result on one retrieval surface will not transfer to another, and every dated observation in this note is a constructed example. Measure your own.
Sources
- Scrunch pricing — Scrunch AI (2026-09-28)
- Profound homepage — Profound (2026-09-28)