Two AI search trackers gave one brand 41% and 23% visibility

Publication date: 18.09.2026
In brief
  • Mike King of iPullRank argues that AI search has no single "true" visibility value to check a tracker against.
  • Model answers are stochastic, so one run of a prompt guarantees nothing and only the number of repeats buys reliability.
  • Two trackers can show 41% and 23% for the same brand, and both can be right in their own way.

The situation Mike King describes on iPullRank will be familiar to anyone who has bought these tools. A client runs two AI visibility platforms side by side. One says the brand shows up in 41% of the prompts that matter, the other says 23%. Someone in marketing has to put a number in a board deck, and the question of which one is correct comes up every week.

±44points is the 95% confidence interval for a tracker that runs a prompt once; fifty runs narrow it to ±14.

Why "accuracy" does not exist here

Classic rank tracking has a reference: the page is in fourth position or it is not. AI answers have none. The model responds stochastically, shaped by user context, model parameters and the moment it called retrieval. King suggests treating it like an opinion poll: nobody asks whether a poll is "accurate", they ask whether the method is stable.

The numbers he cites illustrate it well. In Mike Sonders's study, across 100 ChatGPT runs only about 11% of brands appeared in 80% or more of the responses for a prompt. Earlier research found that 72% of mentioned brands showed up in fewer than one in five responses. Semrush found only 25% domain overlap between ChatGPT's Instant and Thinking modes.

Method choices that break comparisons

A gap between two tools is often explained by different choices their authors made rather than by an error.

DecisionWhat gets measured
API or scrapingModel behavior under controlled conditions versus what a user sees in the interface
Logged out or logged inA clean baseline versus personalization that account history can contaminate
One run or manyAn unreliable yes/no versus a rate with a calculable margin of error

In other words, two trackers are two different instruments measuring different distributions. Comparing their numbers is as pointless as adding Celsius to Fahrenheit. We looked at a similar picture in a piece on AI search metrics worth tracking: without shared counting rules the numbers stop being a metric.

What to ask a vendor and what to do with your own tracking

King advises asking in the demo about the collection method, account state, model choice, geography of queries and what exactly counts as a "mention". Ask for a methodology changelog and raw responses. If you measure in-house, pin the configuration, run enough repeats to detect a real business effect, and tie visibility to referrals, conversions and logs. Change the method mid-year and the trend is broken, so comparing with last quarter no longer works.

For smaller markets there is one more complication the original does not cover. Most trackers collect answers from English-language or US IPs, while your customers ask in Ukrainian and often from a phone. Before a visibility figure goes into a report, find out in which language and from which country it was collected. We wrote about measuring this in practice in our guide to AI search visibility.

I hope the next time someone shows "41%", they are also asked for the confidence interval. My bet is that most vendors will reach for a slide about something else.

SEO Factory Editorial Team

The SEO Factory editorial team reviews primary sources every day — the Google Search Central blog, Search Engine Land, Search Engine Journal and other industry publications — and picks what actually affects search and advertising for businesses in Ukraine.

Every story is checked against the original and comes with a takeaway: what changed and what to do about it as a site owner or marketer.