AI Search Statistics for Ecommerce: 1,800 Answers, 5 Engines, 8 Days
Every AI visibility tool shows merchants a score and a chart. Mine included. I wanted to know whether that number means anything at all.
So I asked five AI engines the same 20 shopping questions about one brand, 18 times, over eight days. 1,800 answers, every one of them kept.
The specific thing I was trying to settle: when a merchant watches their visibility move five points in a week, are they looking at a real change, or the same dice rolled again?
I asked the questions three ways:
- three times inside ten minutes
- five times across a single day
- then once a day for a week
Five engines named 124 different brands as competitors to Gymshark. Gemini alone named 99. Two answers ten minutes apart disagreed almost as much as two a week apart.
Which means a single AI search tells you very little. Here is the proof, and everything I got wrong on the way.
Key AI Search Statistics
of the brands Gemini recommends change between two identical searches. Google AI Overviews changes 27%.
- Google AI Overviews. Swaps 27% of the brands. Reuses 50% of the sources. Features 12 brands.
- Perplexity. Swaps 36%. Reuses 72%, the most of any engine. Features 35 brands.
- Claude. Swaps 52%. Reuses 49%. Features 37 brands.
- ChatGPT. Swaps 55%. Reuses 32%. Features 35 brands.
- Gemini. Swaps 62%, the most. Reuses 30%, the least. Features 99 brands.
Across all five, 124 different brands.
On one engine, a score has to move a long way to mean anything: 15.9 points on ChatGPT, 17.2 Perplexity, 19.3 Claude, 24.8 Google AI Overviews, 28.3 Gemini. Track all five and it drops to 4.1. That is the case for not watching one engine alone.
times a brand appeared in a single answer and was never seen again, from 356 different companies.
how often Gymshark was recommended, depending entirely on what was asked, not which engine was asking.
engines named Gymshark every single time on two questions: influencer recommendations and oversized t-shirts.
of the pages AI cited were product pages. 88% were articles, guides and forum threads.
sources cited in every single run belonged to Perplexity. ChatGPT and Claude had none at all.
citations for Reddit and 1,407 for YouTube, out-citing Gymshark's own site more than two to one.
more citations for a DR 7 site than for Forbes. Five of the 25 permanent sources score below DR 60.
of product questions split into labelled picks like "best with pockets". Brand questions did it 28% of the time.
the range Claude ranked Gymshark across 18 answers to the same question. It named the brand every time.
The Methodology
You cannot separate noise from real change by repeating a question. You need to know when the answers changed.
So I asked the identical 20 questions in three different patterns.
The three tests
3 runs · 10 minutes
Back to back
Three runs inside ten minutes. Nothing in the world can change in ten minutes, so anything that moves here is pure model non-determinism. This is the floor.
5 runs · one day
Across one day
Five runs at two-hour intervals, 9am to 7pm. This catches anything that drifts with time of day, traffic load or index freshness.
7 runs · one week
Once a day for a week
Seven daily runs. This is the cadence a real merchant checks at, and the only test that can show genuine drift rather than sampling noise.
Same 20 questions, same 5 platforms, three timing patterns. 18 runs, 1,800 answers.
Comparing test 3 against tests 1 and 2 is what produces the number that matters: how big does a change need to be before it isn’t just the dice?
I ran test 1 twice, on separate days, because I did not trust the first result. Gemini came out at 0.39 both times, once from two runs and once from twelve.
That is the part of this study I am most confident in.
First, A Control Question
One question was not about recommendations. It checked the engines could hold a fact still: who founded Gymshark, and where is it based. Two checkable facts, on every platform, every run.
Who founded Gymshark: correct in 90 out of 90. Where is it based: correct in 90 out of 90.
Every platform. Every run. Zero variance.
The same engines, in the same breath, could not hold a recommendation still for ten minutes.
AI knows who founded Gymshark. It just can’t decide who its competitors are.
That is the whole study in one line. Not a hallucination problem, and not one better models will fix. A taste problem, and taste gets re-rolled every time.
Recommendation is a ranking problem, and ranking is where the dice live. Everything below is a model choosing again, not a model getting something wrong.
Does Time Between AI Searches Change The Results?
I started with the tightest test I could run. I asked “what are the best gym clothing brands?” three times inside ten minutes.
ChatGPT returned three different brand lists, containing 26 distinct brands between them. Claude returned three different lists containing 19. Perplexity, three lists, 13 brands.
Google AI Overviews returned the identical list all three times.
The worst case anywhere in the data was “where should I buy affordable gym wear online?” on Claude, where three consecutive searches shared just 6% of their recommendations.
The engines are nowhere near equally unstable, which matters more than the volatility itself. “AI recommends X” means nothing without naming the AI.
Then the real question. If AI answers genuinely drift over time, the gap between two answers should widen the longer you leave between them.
It barely does.
| Test | Gap between searches | Brand-set overlap | Brands that differ |
|---|---|---|---|
| Back to back | 10 minutes | 0.54 | 46% |
| Across one day | 2 hours | 0.53 | 47% |
| Once a day | 24 hours | 0.49 | 51% |
Stretching the gap from ten minutes to a full day costs five points of overlap. Roughly nine tenths of the difference between this week’s answer and last week’s was already there ten minutes later.
That is the finding that should change how you read any AI visibility chart. Week-over-week movement is not mostly drift. It is mostly the same dice, rolled again.
But it is not identical across engines:
| Platform | 10 minutes | Same day | 24 hours |
|---|---|---|---|
| Google AI Overviews | 0.77 | 0.80 | 0.63 |
| Perplexity | 0.66 | 0.63 | 0.58 |
| Claude | 0.45 | 0.48 | 0.45 |
| ChatGPT | 0.45 | 0.43 | 0.43 |
| Gemini | 0.37 | 0.32 | 0.37 |
Claude, ChatGPT and Gemini are flat. A week changes their answers no more than ten minutes does. All noise, no drift. Watching them daily tells you nothing.
Google AI Overviews and Perplexity genuinely move. The steadiest engines minute to minute, and the only two that lose real ground over days: AI Overviews 0.77 to 0.63, Perplexity 0.66 to 0.58. On these, a weekly change is worth investigating.
The irony is worth sitting with. The engines steady enough to trust are the only ones where the number actually moves for a reason.
Getting Named In AI Answers Is Reliable. Your Rank Is Not.
This is the most useful pattern in the whole study, and it holds on every platform.
Where a brand owns a question, it owns it reliably.
Gymshark appeared in 27 of 60 question-and-platform combinations in every single run across eight days. That is not a coin flip. That is a position.
But its rank inside those same answers swung wildly:
| Question | Platform | Appeared | Ranked |
|---|---|---|---|
| What are the best gym clothing brands? | Claude | 18/18 runs | #1 to #12 |
| Best gym clothes brands for men | ChatGPT | 16/16 runs | #1 to #11 |
| What gym clothing brands do influencers recommend? | ChatGPT | 16/16 runs | #1 to #9 |
Present in 100% of answers. Ranked anywhere from first to twelfth.
If your tool tells you that you “fell from #2 to #6” this week, it is almost certainly telling you nothing.
Track presence and share of voice. Read rank as a range.
Which AI Engines Are Most Consistent
This is the share of recommended brands rewritten between two consecutive, identical searches:
| Platform | Brands rewritten |
|---|---|
| Google AI Overviews | 27% |
| Perplexity | 36% |
| Claude | 52% |
| ChatGPT | 55% |
| Gemini | 62% |
Google AI Overviews is the most repeatable surface I tested. Gemini is the least. Both are Google, which I still find odd.
Gemini rewrites roughly two-thirds of its recommendations between two identical questions asked minutes apart.
If you are benchmarking your brand on Gemini alone, you are reading tea leaves.
Ghost Competitors in AI Searches
593 times an engine named a brand for a question and never named it again. One appearance in eighteen runs, from 356 different companies. Beyond Yoga. BAM Activewear. Greyson. Janji. P.E. Nation. Hundreds more.
Run a single scan and these look like competitors worth studying.
They are dice rolls. They appeared once in eight days and never came back.
And not only obscure labels. H&M was a one-off on ten question and engine pairs. Reebok on eight. Fabletics on eight. Big brands flicker in and out just as readily.
Every engine does it. The scale differs a lot:
| Platform | Brands named once and never again | Every brand it named | Share that were one-offs |
|---|---|---|---|
| Gemini | 221 | 656 | 34% |
| ChatGPT | 165 | 415 | 40% |
| Claude | 104 | 309 | 34% |
| Perplexity | 65 | 222 | 29% |
| Google AI Overviews | 38 | 134 | 28% |
Gemini invents phantom competitors 5.8x more often than Google AI Overviews. 221 against 38.
As a share it flips. Two in five brands ChatGPT named showed up once and never came back. Gemini looks worse in raw count only because it names more.
Even on the steadiest engine, more than a quarter of the brands in a single scan are noise.
This is precisely why one-shot “AI audit” screenshots mislead, and why any alert that fires on “a new competitor appeared in your category” will be wrong most of the time.
What AI Recommends Instead Of Gymshark
I asked all five engines what competes with Gymshark, 18 times each, expecting this to be the ugliest question in the set. It was.
124 distinct brands were suggested in total. Gemini alone named 99.
The spread between engines on that one question is the widest gap in the study:
| Platform | Brands put forward | Named in every single run |
|---|---|---|
| Gemini | 99 | Lululemon, Nike, Under Armour, Fabletics, Girlfriend Collective |
| Claude | 37 | Nike |
| Perplexity | 35 | AYBL, Alphalete, NVGTN, Vuori, Girlfriend Collective, Rhone |
| ChatGPT | 35 | Vuori, Girlfriend Collective, Under Armour |
| Google AI Overviews | 12 | Alphalete, YoungLA, AYBL, NVGTN |
Gymshark itself is left out of the right-hand column, since the question asked what competes with it.
Ask Gemini who competes with you and it will eventually name almost anyone in the category.
Every engine has a hard core that survives all eighteen runs, and those lists are short. Gemini named 99 brands and 5 showed up every time. Claude named 37 and stuck with one.
Those recurring names are the real competitive set. The other hundred are noise.
A merchant who screenshots one scan and panics about a new rival is reacting to a brand that will not appear again.
On the head-to-head questions, most engines refuse to pick. “Gymshark vs Lululemon” came back a tie in 79 of 90 runs.
Perplexity was the outlier, and consistently so. On the three-way question, “Alo vs Gymshark vs Lululemon”, it picked Lululemon in 12 of 18 runs.
Where AI Citations Come From
Across all 1,800 answers, the most-cited domains were not publishers:
| Domain | Domain Rating | Citations |
|---|---|---|
| reddit.com | 95 | 1,660 |
| youtube.com | 99 | 1,407 |
| gymshark.com | 79 | 668 |
| 7 | 369 | |
| 87 | 354 | |
| 94 | 210 |
Reddit and YouTube out-cite Vogue, GQ and Business Insider combined. They also out-cite the brand’s own website by more than two to one.
Citations churn harder than brands do. 41% of every URL cited appeared in exactly one run.
But against that churn sits a sticky core of 25 sources cited in 100% of runs: three GQ features, two Esquire guides, two Business Insider guides, two Men’s Health round-ups, three Reddit threads, one YouTube video, and single entries from Vogue, the Independent, Men’s Journal and a handful of niche gear blogs.
Those pages decide the answer, every time. Getting onto one of them moves every future answer. A one-off citation moves nothing.
Here is the part I did not expect. Of those 25 permanent sources, 23 belonged to Perplexity. Google AI Overviews had one. Gemini had one.
ChatGPT and Claude had none. Not a single source that survived every run.
Placement durability varies enormously by engine. Perplexity reuses 72% of its sources between asks; ChatGPT reuses 32%.
So a citation is not one thing. On Perplexity you are buying a seat that holds. On ChatGPT you are buying a lottery ticket that gets redrawn every time someone asks.
What Kind Of Page Gets Cited By AI
Almost never a product page. Of the 2,166 distinct URLs cited across the study:
| Page type | Share of all citations |
|---|---|
| Article, guide or forum thread | 88% |
| Category or collection page | 6% |
| Product page | 3% |
| Homepage | 3% |
Narrow it to pages on activewear brands’ own domains, where you would most expect product links, and it still holds:
| Page type on a brand’s own site | Share |
|---|---|
| Article or editorial page | 50% |
| Category or collection page | 21% |
| Homepage | 17% |
| Product page | 12% |
Half of the citations a brand earns on its own domain point at content, not catalogue. Gymshark’s most-cited own page is not a product. It is /pages/we-do-gym, a brand story page.
Optimising individual product pages for AI citation is aiming at the 3%. The pages that get cited are the ones that explain, compare and round up.
Does Domain Authority Matter For AI Citations?
The question I most wanted answered. If AI defaults to the biggest sites, none of this is winnable for a small store. So I pulled the Ahrefs Domain Rating for every cited domain.
Authority clearly helps. The median permanent source scores DR 87, and the household names are all in there: YouTube (99), Reddit (95), Business Insider (92), Vogue (91).
But it is nowhere near a rule.
| Domain Rating | Share of the 25 permanent sources |
|---|---|
| 90 and above | 9 of 25 |
| 80 to 89 | 9 of 25 |
| 60 to 79 | 2 of 25 |
| Under 60 | 5 of 25 |
A fifth of the sources that survived every single run score under DR 60. The lowest, thatfitfriend.com at DR 41, was cited in all 18 runs.
And the sharpest case in the dataset is not close:
That is not a rounding error. It is a site most SEOs would not bother pitching, beating one of the most authoritative domains on the web, in the same answers, on the same questions.
Authority gets you considered. Relevance wins the seat. A tight page answering the exact question beats a big domain that half answers it. That gap is where a small store competes.
You cannot buy your way onto Vogue this quarter. You can earn a place on a DR 40 review blog AI already trusts, and it will hold.
Domain Rating by Ahrefs.
Is Your AI Visibility Chart Lying?
Here is the number this whole study was built to produce.
I compared the spread within a single day against the spread between days. The gap between them is the smallest change you can actually detect.
That is with all five platforms pooled. On a single engine the bar is much higher:
| Tracking | Minimum detectable change |
|---|---|
| All five platforms | 4.1 points |
| ChatGPT only | 15.9 points |
| Perplexity only | 17.2 points |
| Claude only | 19.3 points |
| Google AI Overviews only | 24.8 points |
| Gemini only | 28.3 points |
This is the strongest argument for multi-platform tracking I know of, and I say that as someone who sells one.
Watch one engine and you cannot detect anything short of a 16-point swing. Watch five and the threshold drops to about four.
Breadth is what makes the measurement work at all.
Why The Question Matters More Than The AI Engine
Everything above says the answers move. This is the part that does not.
Gymshark’s presence barely depends on the engine or the day. It depends almost entirely on what the question is about:
All twelve unbranded questions, worst to best:
| Question | Gymshark appeared |
|---|---|
| What brands do fitness influencers recommend? | 100% |
| Best oversized gym t-shirts | 100% |
| What are the best gym clothing brands? | 99% |
| Best athletic wear in the UK | 98% |
| Best gym clothes brands for men | 94% |
| Best squat-proof leggings | 87% |
| Best seamless leggings | 79% |
| Best workout leggings for women | 66% |
| Which activewear brands are actually good quality? | 46% |
| Where should I buy affordable gym wear online? | 45% |
| Best men’s shorts for lifting | 32% |
| Best affordable activewear brands | 26% |
That is a 74-point swing driven purely by the topic, against a noise floor of about four. Which question you own matters roughly 18 times more than when you happen to check.
The Two Questions With No Movement
On influencer recommendations and oversized t-shirts, all five engines named Gymshark in every single run. Not 90%. Every one.
Not random wins. Gymshark was built on influencer marketing, and oversized fits are what it pushes hardest today. The two things the brand is known for are the two the engines never waver on. Own the attribute in the real world and the volatility disappears.
The reverse holds just as hard. Every engine has decided Gymshark is not the affordable option, and no amount of re-asking changes it.
Nobody Wins Every Question
Each question has a different owner, and they are remarkably consistent:
| Question | Owns it | Present in |
|---|---|---|
| Best men’s shorts for lifting | Ten Thousand | 99% |
| Best affordable activewear | CRZ Yoga | 100% |
| Best workout leggings for women | Lululemon | 100% |
| Best oversized gym t-shirts | Gymshark | 100% |
| Best squat-proof leggings | Gymshark | 87% |
| Best athletic wear in the UK | Gymshark 98%, Sweaty Betty 90% | |
| Which brands are actually good quality? | Lululemon 100%, Vuori 93% | |
| Where should I buy affordable gym wear? | nobody reaches 80% |
That is positioning, not luck. Ten Thousand is a training-shorts specialist. CRZ Yoga is a budget label. Sweaty Betty is British. The engines have absorbed what each brand is for, and repeat it.
Two questions had no owner at all. Where nobody has claimed the ground, the answer is pure churn.
How AI Answers Split Into Slots
It goes finer than one brand per question. Engines routinely split a single answer into sub-picks and hand each one to a different brand.
On “best workout leggings for women”, 77% of answers broke themselves into labelled slots:
| Slot inside the answer | Usually won by |
|---|---|
| Best overall | Lululemon |
| Best high-waisted | Lululemon |
| Best with pockets | Gymshark, every time |
| Best moisture-wicking | Gymshark |
| Best for low impact / yoga | Athleta |
| Best budget | Athleta, CRZ Yoga |
Gymshark loses “best overall” on that question. It wins “with pockets” outright.
That is the opening. You do not have to beat Lululemon to get named. You have to own a slot narrow enough that the answer needs someone to fill it.
Segmenting And AI Citations
First, a caveat that turned out to matter. Segmenting only really happens on product-level questions. Ask for a product (“best seamless leggings”) and the answer splits itself 53% of the time. Ask for brands (“best gym clothing brands”) and it splits 28% of the time. Mixing the two together hides what each engine does, so the table below uses product-level questions only:
| Platform | Product answers that split into slots |
|---|---|
| Perplexity | 85% |
| Google AI Overviews | 63% |
| Claude | 53% |
| ChatGPT | 36% |
| Gemini | 30% |
Google AI Overviews is the one to notice. 63% on product questions, 9% on brand questions. Two different behaviours, and averaging them makes it look mid-table when it is second only to Perplexity on the queries that matter.
Slots also track how hard the engine is reading. Answers that segment cite 11.0 sources on average against 7.7 for the ones that don’t. The engines that read more, sort more.
One honest limit, and it is the thing I most wanted to prove and could not. I can show that segmenting engines are the heavy citers. I cannot show that a given slot was lifted from a given page. Of the fifteen answers that named Gymshark “best with pockets”, none cited a page with “pocket” in its title. I captured citation titles and URLs, not the body text of those pages, so a round-up may well carry that section without saying so in its title. Treat the slot-to-source link as a strong lead, not a proven mechanism.
What Merchants Should Do About AI Visibility
- Track presence and share of voice. Not rank. Rank moves for no reason.
- Go for the segmentation. Find the round-ups that already get cited in your category, and get added as the pick for one thing you can actually win. Not “best overall”.
- Get mentioned on sticky citations. The same handful of pages decide the answer run after run, and you will often find easy targets to get featured on.
- Ignore any competitor you have seen once. Most of them never come back.
- Check weekly, or daily. Anything faster is measuring the dice.
Shop Mentions does the tracking part. Every ChatGPT and Google AI Overviews mention of your store, down to product-level prompt tracking.
Then it helps you close the gaps: the sticky citations you are missing, an AI agent you can ask about your own data, and the on-page side too, with AEO articles and AI product optimisations.
Run a free brand scan on your store
The Raw AI Search Data
Everything here is recomputable from the raw responses: 1,800 timestamped answers with full text and citations, the method, the analysis code, and the computed statistics.
Ask and I will send you the lot.
Anyone can write an opinion about AI search. I would rather hand over 1,800 answers and let you check my working.
Method notes
- Dates: 31 July to 7 August 2026. Runs: 18. Answers: 1,800.
- Platforms and models: ChatGPT (
gpt-5-search-api), Claude (claude-haiku-4-5with web search), Gemini (gemini-2.5-flashwith grounding), Perplexity (sonar), Google AI Overviews (via DataForSEO SERP). - Questions: 20 frozen questions in four groups. All presence, rank, churn and ghost statistics use only the 12 unbranded questions. Questions that name the brand echo it by construction and would inflate the numbers.
- Late cells excluded. A handful of answers were retried outside their run’s window after an API failure. All presence, rank, overlap, churn and ghost statistics drop them, so every interval quoted here is the real one. 1,739 of the 1,800 answers carried usable text; 101 fell outside their window.
- Segmentation rates are measured on the seven product-level questions only (specific garments, price points and places to buy). The five brand-level questions ask for a list of brands and rarely segment, so pooling them together flattens the differences between engines.
- One-off brands are counted per question and per platform: a brand that appeared in exactly one of the eighteen runs for a given question on a given engine. The same company can therefore be a one-off on several questions, which is why 593 one-offs come from 356 distinct companies. All 60 question and platform pairs cleared the five-run minimum, so nothing was dropped for a thin sample.
- Limits: one brand, one category, one locale (en-US). The magnitude will differ elsewhere. The direction, I suspect, will not, but one brand is one brand and I would not want anyone quoting these exact percentages as universal.