Measure AI citations without inventing a ranking

A citation report should let another person understand what was observed. That sounds modest, but it is a demanding standard. A percentage without a defined question set or denominator can look precise while saying very little. A screenshot without a date can be difficult to compare with the next run. The first job is to establish an observation method that someone else could follow.
This guide proposes a small working protocol. It is a practical method for organizing your own research, not an industry standard or a claim that a sample represents every user. Start with a manageable panel, record its limits, and keep the underlying answers available for review. A spreadsheet is enough to learn whether the method is useful.
Write the question the report will answer
Before choosing prompts, finish a sentence such as: this report describes which source URLs appeared in responses to these twelve customer questions during this collection window. That sentence defines a narrower and more defensible job than measuring a brand's total AI visibility. It also makes it easier to explain what the report leaves out.
Choose the audience for the report. An editor may care about which research pages are cited. An agency client may care about examples related to a particular product category. A product team may be comparing two permitted collection methods. Those purposes lead to different query panels. Combining them without labels can make a report harder to interpret.
Build a panel from actual reader tasks
List questions that correspond to specific needs: comparing an option, checking a requirement, troubleshooting a problem, or locating a reference. Record where each question came from, such as a customer interview or an existing support topic. Do not describe the panel as representative unless you have a sampling method that supports that claim.
Keep branded and unbranded questions in separate groups. A question containing a company's name creates a different context from a general category question. Also separate broad questions from narrow ones. The point is not to make every group the same size, but to avoid averaging unlike observations into a number whose meaning no one can explain.
For a first exercise, select twelve questions in three groups of four. Freeze the wording for one reporting period. If a question is later improved, assign the revision a new version and note when it entered the panel. Quietly editing prompts midway through collection can make an apparent change difficult to interpret.
Define one run before collecting anything
One run should identify the question, interface, collection time, and relevant context. Record the provider and product surface. If you use an API, record the model identifier and options that affect the request. If you use a browser interface, record the visible product setting and any relevant account or location context you can reliably observe.
Avoid placing private account details in a shared report. A simple context label can preserve the distinction without exposing credentials or personal information. Record whether the session began fresh or continued an earlier conversation. Previous messages can be part of the question's context, so a follow-up answer should not be silently compared with a fresh-session answer.
Decide the number of repeats in advance. Three runs per question can be a practical starting exercise, but it is not a magic threshold for statistical confidence. Keep the schedule consistent enough for your purpose and explain why it was chosen. If a service is unavailable, record the failed attempt separately from a completed answer with no citations.
Preserve the answer and the references together
Save the exact output you are permitted to retain, its cited URLs, and a collection timestamp. Keep an unedited version before adding your own classification. If you normalize URLs for reporting, retain the original URL too. Redirects and tracking parameters can make duplicate handling useful, but normalization should not erase evidence that another reviewer may need.
OpenAI's web search guide documents citation annotations with a URL, title, and location in the answer text. Those fields illustrate why source and answer records belong together. Use the fields actually available in your collection environment. Do not invent a missing citation position or assume another interface provides the same response structure.
It is useful to distinguish a cited source from a brand mention. An answer might mention a company while linking to a third-party article, or cite a company page without naming the company in the prose. Keep those observations in separate columns. A client can then understand which question each measure answers.
Choose the denominator before showing a percentage
Suppose the twelve-question panel is repeated three times, producing thirty-six scheduled runs. Two attempts fail, and thirty-four answers complete. Eight completed answers cite at least one URL from a selected domain. You can report eight of thirty-four completed answers, alongside the two failed attempts and the full schedule. This is a hypothetical arithmetic example, not a performance benchmark.
A different measure counts individual citation occurrences. If one answer cites the same domain twice, it contributes two occurrences but only one answer with a citation. Both measures can be useful, but they answer different questions. Label them in ordinary language rather than giving both the same visibility score.
If you restrict a measure to answers that contain any citations, say so. Eight out of twenty citation-bearing answers is a different denominator from eight out of thirty-four completed answers. Show the counts beside percentages. Small samples are easier to understand when readers can see the actual number of observations behind a chart.
Inspect changes before explaining them
When the next period looks different, check the collection records first. Did the query wording change? Was a different interface used? Did the source URL redirect? Were there more failed runs? A difference may still be worth investigating, but the report should separate the observed change from possible explanations.
Google's AI features documentation notes that its AI experiences can show different responses and links. Treat each defined collection environment as its own series. A pooled figure across unrelated surfaces may obscure the very differences the report should help a reader examine.
An annotation can say that a publisher updated a page between periods. It should not say that the update caused a citation increase merely because the events occurred in that order. To explore causation would require a more deliberate study, with attention to other changes and the limits of the design. Most small operational reports are better used to identify questions for follow-up.
Give the reader a short audit trail
A useful report ends with the panel version, collection window, completed and failed run counts, and a link to the permitted supporting records. Include one example that illustrates the main observation and one limitation that materially affects interpretation. Avoid burying the method in a separate document that the reader cannot access.
Before sharing, ask a colleague to reproduce one row from the saved evidence. Can they find the prompt, answer, source link, and classification reason? If not, repair the record before adding more metrics. The result should be a report that invites inspection and helps someone decide what to investigate next.