---
title: "How to measure the AI agent channel properly · Specoria"
description: "Asked the same question a hundred times, AI assistants almost never gave the same list. Measure it with real buyer tasks, repetition and confidence intervals."
source_url: "https://specoria.com/insights/measuring-the-agent-channel/"
lang: "en"
---

[← All insights](https://specoria.com/insights/)
Measurement October 6, 2026 7 min read

# Measuring the agent channel: why a single screenshot is misleading

In a study that asked the same question a hundred times, AI assistants almost never gave the same list. Measuring the agent channel properly takes real buyer tasks, repetition and confidence intervals, not rankings.

[Specoria Team](https://specoria.com/about/)

A marketing manager asks ChatGPT, "What are the best espresso machine brands?" Their brand isn't on the list. The screenshot ends up in a meeting. The next day, a colleague asks the same question, and this time the brand is third. Which one is right?

The answer: on its own, neither tells you anything. In this article, we explain why AI assistants' answers vary, what can be measured reliably despite that variation, and how to tell whether a fix actually worked.

## Same question, a different list every time

Rand Fishkin of SparkToro and Patrick O'Donnell of Gumshoe tested this systematically in research published in January 2026. Over November and December 2025, 600 volunteers ran 12 different recommendation prompts through ChatGPT, Claude and Google's AI answers a total of 2,961 times. The prompts covered categories ranging from chef's knives and headphones to hospitals and digital marketing consultants.

The results are clear:

- Ask ChatGPT or Google's AI the same question a hundred times, and the chance that any two answers give the **same list of brands** is less than 1%.
- The chance of two lists coming back **in the same order** is roughly one in a thousand.
- By contrast, **how often** certain brands appear in answers is far more consistent. For the question about headphones for family travel, the leading brands show up in 90% to 100% of answers.

The researchers' conclusion is the ground rule for anyone measuring the agent channel: in AI, "ranking position" isn't a meaningful metric, but the **rate at which a brand appears across many answers** is.

## Where does the variation come from?

- **Probabilistic generation.** Language models generate every answer based on probabilities. The same question can produce a different answer, even on the same model.
- **Shifting sources.** Assistants that search the web can find different pages each time and rely on different sources.
- **Model updates.** Models and search infrastructure are updated frequently. Last month's measurement may not reflect today's model.
- **Context.** Language, location, the wording of the question and the user's previous conversations all affect the result.

That's why a single answer, good or bad, is an anecdote. Measurement only starts when you look at the distribution across many answers.

## What to measure: selection rate

In agentic commerce, the real question isn't "Were we mentioned?" It's "**Did we make the shortlist?**" That's why we measure selection rate: for a given buyer task, the percentage of answers in which your brand appears on the recommended shortlist.

But the rate alone isn't enough. How many answers it's based on matters too. An example:

- You were selected in 6 of 20 answers: a rate of 30%. But the 95% confidence interval (Wilson) runs from roughly **15% to 52%**. In other words, the true rate could mean "rarely" or "about every other answer."
- You were selected in 24 of 80 answers: the rate is still 30%, but the interval narrows to **21% to 41%**.

The same rate carries very different certainty at different sample sizes. An AI visibility score reported without a confidence interval is a score that doesn't tell you how much to trust it.

## Six rules for a sound measurement setup

- **Buyer tasks, not keywords.** "Espresso machine" is a search term. "Find an espresso machine under €500 that ships in two days and comes with 30-day returns" is a buyer task. Agents filter by constraints, and only constrained tasks show where your store gets dropped.
- **Repetition.** Run every task many times, on every surface and in every language, not just once. We treat any cell (one task on one surface in one language) with fewer than 20 answers as an "insufficient sample."
- **Confidence intervals.** Report every rate with its interval. Don't make decisions based on a result with a wide interval.
- **Keep surfaces and languages separate.** ChatGPT, Gemini, Claude and Perplexity rely on different sources and different models. Questions in Turkish and in English also give different results. Rolling them all into one "AI score" hides the place that needs fixing.
- **Flag version changes.** When the model, the prompt template or the evaluation method changes, draw a break line in the measurement series, and don't compare the two sides of that line.
- **Record the reason.** Alongside "Are we on the list?", record why a competitor was picked and which sources the answer cited. That's where your fix list comes from.

## What web analytics shows, and what it doesn't

According to OpenAI's guidance for publishers, ChatGPT automatically adds the `utm_source=chatgpt.com` parameter to the links it includes in search results. That lets visits from ChatGPT show up as a separate source in analytics tools. In your server logs, you can also track requests from OAI-SearchBot, OpenAI's search crawler, and from ChatGPT-User, which opens pages on a user's behalf.

These are valuable signals, but they only measure the **visible side**: the conversations where you were recommended and someone clicked. Conversations where you didn't make the list leave no trace in analytics. There's no visit and no bounce. Selection rate measures exactly that invisible side. Together, they give you the full picture: one shows how much traffic the channel brings in, the other shows how much you're missing.

## How do you know a fix worked?

You've fixed your product data and rewritten your return policy. To see the effect:

- **Measure first.** Before the fix, build a baseline of at least a few weeks on the same tasks.
- **Log the date.** Note the day the change went live. Search systems can take time to pick up new content, so leave the first week out of the comparison.
- **Repeat under the same conditions.** Same tasks, same surfaces, same languages, same prompt template.
- **Run a statistical test.** Compare the two periods' rates with a two-proportion test.

An example shows why the test matters. Your rate went from 30% to 45% on 20 answers. A 15-point gain looks good, but even if nothing had actually changed, the chance of a gap this size appearing at random is about one in three (p ≈ 0.33). So the honest call is "no change." If the rate went from 20% to 40% on 80 answers each, that same probability drops to about six in a thousand (p ≈ 0.006). That's an improvement chance can't explain.

Keep tracking the tasks you didn't touch, too. If they rose over the same period, the change may come from the model itself, not from your fix.

## The bottom line

In the agent channel, a screenshot is an anecdote; a measurement is a distribution. Sound measurement starts with real buyer tasks, repeats every cell enough times, reports every rate with a confidence interval, and confirms the effect of a fix with a test. Without that discipline, "our AI visibility went up" is just a hope.

Specoria works exactly this way: it measures your store with real buyer tasks across four surfaces and multiple languages, traces the points where you get dropped back to their source, and re-measures the impact of each fix with the same tasks. To see your store's selection rate today, request a [free readiness report](https://specoria.com/#contact).

### Sources

- SparkToro, [NEW Research: AIs are highly inconsistent when recommending brands or products; marketers should take care when tracking AI visibility](https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/), Rand Fishkin and Patrick O'Donnell, January 28, 2026.
- Search Engine Journal, [AI Recommendations Change With Nearly Every Query: SparkToro](https://www.searchenginejournal.com/ai-recommendations-change-with-nearly-every-query-sparktoro/566242/), January 30, 2026.
- OpenAI Help Center, [Publishers and developers FAQ](https://help.openai.com/en/articles/12627856-publishers-and-developers-faq).
- OpenAI, [Overview of OpenAI Crawlers](https://developers.openai.com/api/docs/bots).
- E. B. Wilson, *Probable Inference, the Law of Succession, and Statistical Inference*, Journal of the American Statistical Association, 22(158), 1927.

## Related

- [Shopify opened checkout to browser agents. What WebMCP means for every other store](https://specoria.com/insights/shopify-opens-checkout-to-browser-agents-webmcp/)
- [The Turkish checkout through an AI agent's eyes: seven obstacles and how to fix them](https://specoria.com/insights/turkish-checkout-through-an-ai-agents-eyes/)
- [What tires a shopping agent: 33 frictions that slow AI agents down in your store](https://specoria.com/insights/what-tires-a-shopping-agent/)
- [Agentic commerce statistics](https://specoria.com/agentic-commerce-statistics/) Sourced, dated figures you can cite
- [Agent Readiness Index 2027](https://specoria.com/index/) How the benchmarked stores scored, with open data

FREE TEST

## Is your store ready for shopping agents?

See in seconds how a shopping agent reads your store: free, instant, no email needed.

The free test runs in your browser and needs JavaScript. [Request the free deep audit by email instead →](https://specoria.com/#contact)

Your result opens on its own page. Want the full picture? The free deep audit and panel come next, by email.
