When an AI agent recommends, one brand usually leads. When it has to pick one product and act, that brand usually holds, but not reliably. In 25% of the scenarios we tested, a different brand won once the agent acted.
TL;DR
- Measure being picked, not only being named. In 25% of the scenarios we tested (5 of 20), a different brand won once the agent acted.
- Expect the change where no brand clearly leads. Three of the five scenarios that changed hands were in footwear. The six brands that won every recommend run kept their lead.
Most brands preparing for AI shopping are working on getting recommended. The bet is that if AI names your brand, an agent will buy your brand when a shopper hands it the job. We tested that bet in the KNWN Agentic Purchase Study, and it does not always hold.
We ran the study on Codex, OpenAI's agent, with web search and a browser, and gave it the same 20 shopper requests 400 times. In 133 runs it recommended products. In 267 it acted on its own: it chose one product, or chose one and added it to a cart. Nothing was purchased.
The leader usually held. When it didn't, the fall was steep. One brand that led more than half of the agent's recommendations in its scenario was picked in just 2 of 18 runs once the agent acted.
Most leaders held, but five scenarios changed hands
Being an agent's favorite recommendation gives a brand a head start when the agent acts, not a guarantee.
This happens because the agent does a different job when it acts. It narrows to one product it can confirm, and the brand that was easiest to recommend is not always the easiest to confirm. Parts 2 to 4 of this series show how that works.
In each scenario, we found the brand the agent recommended most often. It did not lead every recommend run, because the agent's answer varies from run to run. Across the 18 scenarios with a single most-recommended brand, it led 71% of recommend runs (87 of 123) and was picked in 60% of acting runs (149 of 247). Scenario by scenario, the most-recommended brand was also the brand picked most often in 12 of 20. A different brand won in 5, and 3 had a tie on one side.

Where the winner changed, it changed by a lot. In everyday sneakers, New Balance was the top pick in 67% of recommend runs and the pick in 22% of acting runs. In power banks, INIU went from 71% to 21%, and Anker was picked most. In coffee grinders, SHARDOR went from 56% to 11%, and Breville was picked most.

For a brand, a strong place in AI recommendations protects you most of the time. It does not tell you when it won't.
The changes happen where no brand clearly leads
An overwhelming recommendation lead held when the agent acted. A narrower lead often did not.
When one brand is the clear answer to a need, the agent finds it, confirms it and moves on. When several brands are close, the pick depends on which one the agent can reach and confirm in that run. In our runs, most brands that lost were never compared on their merits at all, which Part 3 covers.
Six brands won every recommend run in their scenario: Soundcore headphones, COSORI air fryers, Oral-B toothbrushes, Laifen hair dryers, Travelpro suitcases and TOMTOC backpacks. All six kept their lead when the agent acted and were picked in at least 5 of 6 acting runs. The five brands that lost the pick had smaller leads, with top-pick shares between 50% and 71% when recommending. Three of the five were in footwear: road-running shoes, everyday sneakers and stability running shoes.
If your category has a clear leader and it is you, recommendations are a fair guide to what an agent will pick. If your category is contested, they are a weak guide, and the pick can go to someone else.
One scenario up close: everyday sneakers
The sneaker scenario shows the gap at its sharpest: one brand led the recommendations and another won the picks.
The request was men's US size 10, standard width, lightweight sneakers for everyday walking, under $120. When recommending, the agent named New Balance as its top pick in 6 of 9 runs, Skechers in 1, and gave no clear pick in 2. When acting, it picked Skechers in 13 of 18 runs, New Balance in 4 and Allbirds in 1.

The run logs show why. In the choose runs we read closely, New Balance was never rejected because the agent judged the shoe worse. In four of the Skechers wins, the agent considered New Balance but could not confirm that the required option was available. In the other three, New Balance never reached the final set. In the two runs where the agent confirmed the New Balance Fresh Foam Arishi was in stock at Zappos, it chose New Balance both times.
New Balance lost those picks on availability the agent could not confirm. That is something a brand can check on its own pages and on the retail listings an agent reads.
What does this mean for ecommerce teams?
Being named and being picked are two numbers. Track both.
AI visibility tools count how often a brand is named in answers. That is the recommendation number. The second number is how often an agent picks you when it has to act on a shopper's behalf. In our runs the two agreed in most scenarios and split in contested ones, where it matters most.
- Count across many runs. Ask the same question repeatedly and record the share of runs that name you, not a single answer.
- Ask both ways. Ask an AI agent to recommend, and separately ask it to pick one product to buy. Compare the two shares.
- Look hardest at contested categories. Where no brand clearly leads, check that an agent can confirm your exact variant, current price and stock on the pages it reads.
Start by asking an AI agent to pick one product in your category ten times, and count how often it picks you.
Next week, Part 2 looks at how an agent shops differently when it acts: it stops comparing and starts confirming.
Get the full report
The full report covers all 20 scenarios and the method. Read the report at knwn.app/benchmark. If you want to see how an AI agent handles your store, talk to us.
Interpretation note: this article is educational material, not a product commitment or a guarantee of rankings, citations, traffic or commercial outcomes.