Every AI visibility tracker turns AI answers into a score: a visibility percentage, a share of voice, a "rank". Two pieces of research explain why those numbers move even when nothing about your brand has changed.
AI answers change almost every time you ask
In January 2026, SparkToro published a study run with Gumshoe (which sells an AI tracking tool, a conflict the authors disclose). 600 volunteers asked ChatGPT, Claude and Google's AI the same 12 questions, 2,961 times in all.
- ChatGPT and Google's AI gave the same list of brands in fewer than 1 in 100 pairs of answers; the same list in the same order, about 1 in 1,000.
- What was stable was how often a brand appeared across many runs. One hospital showed up in 69 of 71 ChatGPT answers, but was named first in only 25.
- The authors' advice: ask the same question something like 60-100 times and average.
Gumshoe's own follow-up works through the statistics: about 100 answers to know a mention rate within ±10 points, about 400 for ±5 points, for one prompt on one engine.
The API isn't the app
Most trackers query AI models through their official APIs, because scraping the consumer apps is slower and riskier. In June 2026, Brandlight (which sells app-based monitoring, so this result favours it) asked 900 questions both ways:
- The brands named in an API answer overlapped with the brands in the matching consumer-app answer by 6-12% on average, depending on the engine.
- The first brand named matched only 9-16% of the time.
- The ChatGPT app read about 3.5 times as many sources per answer as the API.
The two sets were collected a day apart and Brandlight has a commercial interest, so treat the exact figures with care. The direction is consistent with the vendors' own warnings: Semrush's free checker says you may see different answers when you test prompts yourself, and Frase calls its free result a snapshot.
What to look for in a tracker
| Question to ask | Why |
|---|---|
| How many times is each prompt run per period? | One run is a snapshot; dozens give a usable mention rate. |
| App or API? | API answers can differ a lot from what customers see. Few vendors say which they use. |
| Does it show a margin of error? | A visibility score without one can't tell a real change from noise. |
| Mention rate or "position"? | Position is the least stable number in the research above. |
Prices and engines per tool are in our price index. We'll publish our own accuracy test, the same prompts in every tool compared with many real runs, as soon as it's done. How we'll test.
Sources (all read Oct 8, 2026)
- Rand Fishkin, "NEW Research: AIs are highly inconsistent when recommending brands or products", SparkToro, Jan 2026 (study with Gumshoe)
- Search Engine Journal, "AI Recommendations Change With Nearly Every Query: SparkToro", Jan 30, 2026
- Gumshoe (vendor), "How Much Data Do You Need to Measure AI Visibility with Confidence?", Feb 11, 2026
- Brandlight Research Lab (vendor), "The UI vs API Gap: AI Measurement Study", June 2026: research.brandlight.ai
- Semrush AI Search Visibility Checker FAQ; Frase AI Visibility Checker FAQ (vendor pages)