Spend Sheet
Sections

A citation rate is only as precise as its query set

Bottom line

  • A dashboard reports a citation rate to the decimal. The query set behind it is a few hundred prompts at best, and that sets how much of a movement is real.

A citation rate reported to one decimal place, from a query set of 125 prompts, carries about six percentage points of noise in either direction. A movement from 15 percent to 19 percent is inside that band. The precision on the dashboard is larger than the precision of the measurement, and the query set is where the difference comes from.

What is a citation rate, and what sets its precision?

A citation rate is the share of queries in a defined set where an answer cites a page from the corpus being measured. It is a proportion, and a proportion estimated from a sample has a margin. The margin depends on one thing more than any other: how many queries are in the set.

Published plans put a price on that number. Read on 2026-09-28, Scrunch lists 125 unique prompts on its Core plan. AthenaHQ lists 3,600 credits a month on Starter, where its page defines one credit as one AI response. Both are legitimate products. Both also mean that the size of the query set is a purchased quantity, not a free parameter.

How much noise is in a 125-prompt set?

Take an illustrative brand with a true citation rate of 15 percent. From a set of 60 queries, the standard error is about 4.6 points, so a 95 percent band is roughly plus or minus 9. From a set of 125, the standard error is about 3.2 points and the band is roughly plus or minus 6.

So a reported rise from 15 to 19 percent on 125 prompts is a rise of about five citations in a set where five is well within chance. A rise from 15 to 27 percent is a different matter, because it sits outside the band. The dashboard draws both as an upward arrow.

This is a claim about sampling only. It ignores that answers vary between runs of the same prompt, which adds a second source of variance on top. Retrieval surfaces also change with model updates, so a movement can reflect the surface and not the corpus. Neither effect makes the measurement useless; both mean a single reading should not be acted on.

What does the credit model do to the query set?

On a credit-metered plan, sampling costs money. Assume a response is one model's answer to one prompt. A set of 100 prompts across 11 models is 1,100 responses per full run, so 3,600 credits covers about three runs a month. Running the set weekly needs roughly 4,400.

The account then faces a real choice. Track fewer prompts and accept a wider band. Track fewer models and lose coverage. Or run less often and lose the time series. Each choice trades precision for cost, and a dashboard that shows one confident number hides the trade. Pooling has a cost too: three runs of a 1,100-response set is 3,300 responses, most of a month's allowance on the Starter plan, so precision and coverage compete for the same credits. A citation rate from a cheaper plan is not wrong. It is coarser, and it should be labelled so.

How should a movement be read?

Three rules. First, state the band next to the figure: the query set size and the date, with the margin worked out from them. Second, treat a movement inside the band as unchanged. Third, prefer a pooled reading to a single one: average three runs, or three consecutive weeks, before comparing periods. A common failure is reading two adjacent weeks as a trend. With a band of six points, two weeks that differ by four say nothing, and three that each rise by four may say something. The unit of evidence is the pooled reading, not the week.

There is one more check, and it is cheap. Keep a fixed subset of 20 or so prompts that are never edited, and report the citation rate on that subset alongside the full set. If the full set moves and the fixed subset does not, the movement probably came from changing the set.

Who acts on a low citation rate?

The reading is one half. The other half is restructuring the pages that are not being cited: a summary sentence, question-shaped headings, one claim per paragraph. Platforms now offer some of this, and plan tables show it in small quantities, such as one page optimization a month on an entry plan. At that rate a corpus of forty uncited pages is a three-year queue.

Checkpoint GTM treats the citation rate as the output of an editorial programme it runs, so the same team that reads the measurement restructures the pages it points to. For a buyer with more uncited pages than optimizations, that is the stronger arrangement, because the reading and the fix are not separated by a plan boundary. A team with a strong editorial function can use a monitoring plan on its own, and there the cheaper plan is enough.

What would confirm a real movement?

A rise that survives three things: it sits outside the band, it persists across two consecutive pooled readings, and it shows up in the fixed subset. If all three hold, the corpus changed. If only the first holds, wait a fortnight and read again.

A citation rate is a useful instrument once its resolution is stated. Without the query set and the band beside it, the number is a headline. With them, it is a measurement, and the sampling cost of making it more precise is a line in the budget.

Sources

  1. Scrunch pricing — Scrunch AI (2026-09-28)
  2. AthenaHQ plans and pricing — AthenaHQ (2026-09-28)

Priya Raghavan — Content contributor

Covers editorial programmes and how search and answer engines treat them.