TL;DR

Digiday’s look inside the scramble to measure AI visibility is a good snapshot of where the industry is. An IAB survey of 200 brand and agency decision-makers found 45% name measurement as their biggest challenge when comparing AI and traditional buying journeys. The IAB is now building a measurement framework. Its VP of AI explained why this is hard:

“Models aren’t deterministic… the market is messy.”

— Caroline Gigerisch, VP of AI, IAB, quoted in Digiday

That’s true. It’s also usually where the conversation stops. “Not deterministic” can mean the answer barely moves, or that it’s a coin flip. Those call for very different reports. So we measured it.

The test

In one tracking round in August, we asked the core buying questions for five businesses three times each, on each of four engines: ChatGPT, Claude, Gemini and Perplexity. The businesses were a wealth manager, a fractional CFO firm, an SEO agency and two cannabis dispensaries. Same question, same engine, same day. That gave 176 question-and-engine pairs, each with three answers. For each pair we asked two things. Did the business appear all three times? Did the first-listed brand stay the same?

Presence is fairly stable

Did the business appear across three identical asks? (176 pairs)
OutcomePairsShare
Appeared all three times8045%
Never appeared7945%
Appeared some times, not others1710%

Nine times out of ten, the engine gave the same verdict on the business every time. The flips cluster in a predictable place: questions where the business is a borderline answer. Of the 97 pairs where the business showed up at least once, 17 were unreliable. That’s 18% of what you’d call a “win” from a single ask.

Engines differed. Claude flipped on 1 of its 29 pairs. Gemini flipped on 6 of 49, and ChatGPT and Perplexity on 5 of 49 each.

Rankings are not stable

The first-listed brand is a different story. Across the 145 pairs where all three answers listed at least one brand, the top brand stayed the same in 78 and changed in 67. That’s 46%. Ask twice, and nearly half the time a different business leads the list.

This is the number that matters for anyone selling or buying “AI rank tracking.” A screenshot saying “you’re #1 in ChatGPT for best advisor in San Diego” is one draw from a distribution. The next buyer to ask may see someone else on top.

How to measure a system that won’t sit still

Randomness makes a single observation meaningless. Every rule we use comes from that one idea:

  1. Freeze the questions. Use the same question set every round. If the questions change, you can’t tell whether the answers moved or the test did.
  2. Measure rates, not moments. Report the share of answers naming you across dozens of questions and several AI tools. One answer tells you almost nothing.
  3. Repeat the questions that matter most. Asking core questions more than once shows how much of a change is just noise.
  4. Show what every number is out of. A move smaller than the normal spread is not a trend, and the report should say so.
  5. Treat position as an average. “Usually second or third” is honest. “Ranked #2” is not.
  6. Know when the instrument changed. Providers swap the model behind a name without notice. Record the model each answer came from, and flag the round when it changes.
  7. Missing isn’t zero. An engine that failed to answer isn’t an engine that ignored you. Leave it out of the math and say so.

As The Now Agency’s Gabe Feldman told Digiday, “There is no silver bullet in any of this.” None of this is exotic, though. It’s how survey research has handled noisy respondents for decades.

Limits of our data. This is five businesses, one round, three asks per pair. A larger test would tighten the numbers but is unlikely to flip the pattern of stable presence and unstable order. Answers were read by a language model against a fixed rubric. Any bias in that reading applies equally to all three asks, so it doesn’t create the flips.

Frequently asked questions

Do AI assistants give the same answer every time?

Not exactly. In our test of 176 question-and-engine pairs asked three times each, whether a business was mentioned changed in 10% of pairs, but the first-listed brand changed in 46%.

Is AI rank tracking reliable?

A single ranking is one draw. In our data the top-listed brand changed between identical asks nearly half the time. Report average position and presence rates across many questions instead.

How many times should you ask each question?

Enough to see the spread. Our Full Audit asks every buyer need two or three ways, asks one question in ten a second time, and shows every number with what it’s out of. For a single important question, more asks give a steadier estimate.

Want a real audit?

Find out where you reliably show up in AI answers.

BeCited asks your buyers’ questions more than one way and never counts a missing answer as a “no.” See what ChatGPT and Google’s AI tell buyers in your city about your business. The $20 Snapshot asks about 25 of their questions.

See where you stand: $20 Read the methodology

Sources. IAB survey figures and quotes from Caroline Gigerisch and Gabe Feldman are as reported by Seb Joseph, Digiday, September 15, 2026. Stability figures are BeCited tracking data from the August 2026 round, in which core questions were asked three times per engine.