The short version
- A score that moves moderately between reports is expected AI-assistant behaviour.
- The fix is a wider check, not a more frequent one. A score built on hundreds of real buyer questions is stable enough to compare between reports; one built on 15 prompts is not.
- When something real does change, the source list shows it before the score does.
Why does your AI visibility score keep moving?
Moderate movement between reports is the expected behaviour of AI assistants. Nothing about your brand has to change for the number to move.
Evertune, which studies how models represent brands, describes the mechanism plainly: a model predicts likely text rather than looking up a fact, so two identical prompts can land on two different sets of brand names. Add live web search and the pool of things it can read shifts hour to hour.
So a score that slips a few points between runs is not evidence the assistants changed their mind about you. It is one sample, next to another sample.
This is uncomfortable, because it removes the reading most tools are built to give you.
What makes a check stable enough to compare?
What makes a check stable is the number of questions behind it. A wide check averages out the variation. A narrow one is mostly variation.
We looked at eight Adacity reports across eight unrelated direct-to-consumer brands in six categories, 2,156 buyer conversations in total. We sorted every question by one thing: did the buyer type the brand's name, or not?
When the name was in the question, the assistants named the brand back in 98.9% of answers. When the name was not in the question, they named it in 11.2%. The same brands, the same models, the same week. An 8.8x gap, produced entirely by how the question was phrased.
That gap is why the shape of your prompt list decides your score before the models do. Fifteen prompts you wrote about yourself will return a high, stable, meaningless number. It will look reassuring for months while measuring only your phrasing. The way out is to start from the questions buyers ask rather than the ones you would type, which we went into in who is actually asking AI about your brand.
So when a check disagrees with the one you ran last month, the first question is not what changed at OpenAI. It is whether either check asked enough of the right questions to be worth comparing.
What actually changes, when something changes?
When something real shifts, it shows in one of three places, and none of them is the score.
Your sources. A publication that used to describe you stops updating, or a new roundup starts carrying you. What the assistants repeat about you follows what they can read about you.
Your competitors' sources. Someone else earns a place in the comparison articles for your category. Your own coverage did not move; the shelf around you did.
The category's shape. A category with no settled answer converges on two or three names, or a settled one cracks open. This is the slowest and the most consequential.
All three are changes in the world, not in the wording of an answer. They also take weeks or months.
Where does a real change show up first?
A real change shows up first in the list of sites the assistants read to answer, well before the score settles anywhere new.
AthenaHQ, which studied what gets cited, found top-decile content cited 87.0% of the time against 38.6% for the bottom decile. A small number of sources carry most of the answer in any category. When one of them changes what it says about you, that is visible immediately, and it is visible as a specific page you can go and read.
The score moves later, if at all, and by then you have lost the explanation. A number tells you something is different. The source list tells you what changed, and where.
These sources are mostly not the ones you already track. ConvertMate's 2026 benchmark, citing BrightEdge, found 83% of the pages the assistants cite come from outside the organic top 10. Your rankings dashboard is watching a different set of pages than the assistant is reading. We went further into what to do with a citation list in what are AI citations actually good for.
So how often should you re-check?
For most companies, checking monthly or quarterly is sufficient.
A useful cadence follows how fast the answers actually change. Your sources and your category move on the order of months; a weekly read adds false alarms, not information. Re-check sooner when something in the world actually happened: you launched, a competitor did, a major publication covered your category, you rebuilt the pages the assistants read.
The exception is scale. If you sell hundreds of products and a single line disappearing from AI answers costs real money, continuous monitoring earns its price, and several tools do it well. We compared what each of them tracks in the best AEO tools. But a small prompt list still swings between runs, no matter how often you check it.
Ask the questions your buyers actually ask, in their words, across the assistants they use, and count only the times you come up when your name was not in the question. Then keep the source list. That is the thing you will compare against next time.
Want a baseline worth comparing to? Run the check. It is free to start. The $39 full-depth run asks 100+ questions across ChatGPT, Gemini, and Claude, scores only the answers where an assistant raised you on its own, and traces every source behind them, with no subscription required.