How we measure citations.
AI answers are non-deterministic, personalized, and change over time. A citation report that hides that eventually reads as vibes, the exact thing we sell against. So here’s exactly how the audit and the monthly report are made, real models and all, including where it falls short.
We query the real models, not a simulation
Every question goes to the five answer surfaces most people actually use, with live web search turned on. These are the real models’ answers, measured, not our best guess at what they’d say.
OpenAI's current ChatGPT model, queried live with web search on through the OpenRouter API. The same generation your customers are actually using, not a budget stand-in.
Google's Gemini model, queried live with web search on through OpenRouter.
Anthropic's model, queried live with web search on through OpenRouter.
Perplexity's own live retrieval, the Pro model, queried through OpenRouter.
Google's own model answering with live Google Search grounding, the same stack behind the AI Overview at the top of a Google results page.
The honest caveat: the web-search grounding on the ChatGPT, Gemini, and Claude calls approximates, but is not identical to, each product’s own consumer app. Perplexity Sonar Pro is Perplexity’s real retrieval, and the Google AI Overviews line is Google’s own model answering off live Google Search. We report per engine so you see where you stand in each, not one blended number that hides the differences.
The 8 to 10 questions a local actually asks
Not “is my SEO good.” The real research questions a customer types, tuned to your category and your area. A slice of a real set:
- best breweries near [your town]
- date-night winery with a view near [your town]
- where can I try mead near [your town]
- dog-friendly taproom near [your town]
- what's on tap at [you] right now
- does [you] take reservations
- can I get [you] shipped to my state
- best cidery in [your region]
We hold the set steady month to month, so movement in your report is real and not an artifact of changing the questions.
Where we ask from, and what counts
Location
The models can’t see where a searcher is, so “near me” would answer for anywhere. We write your town into every question instead, so your report reflects how the answer looks to someone searching in your area, not a national average.
What counts as a citation
Your business named in the answer text, not merely a link in a sidebar. For each question we record one of three: named, not named, or named with an error, wrong hours, wrong category, closed when you’re open. That third bucket matters as much as the first.
What this can’t tell you
Said plainly, because the honesty is the point.
A single run is a snapshot, not a verdict. Ask the same engine the same question twice and you can get two answers. We read the trend across months, not one result.
Personalization we can’t see. A real person’s history and settings shape their answer in ways our run can’t reproduce. We measure the baseline, not every individual’s screen.
No guarantees. Citations move with the whole web, reviews, lists, and the models themselves. Nobody honest can promise a specific one. We move the signals we can and report what actually changed.
See your own scorecard.
The free audit runs this exact method once, so you can see where you stand before deciding anything. Ongoing monitoring is the Watch plan, $249 a month.
Get my free audit