Key takeaways

  • Researchers at the Wharton School at the University of Pennsylvania tested how consistently AI shopping agents recommend products when the…
  • They used the ACES simulator (Agentic e-Commerce Simulator), which shows the AI agent a screenshot of a product page.
  • A stable decision process should produce the same result given identical content.

What happened

Researchers at the Wharton School at the University of Pennsylvania tested how consistently AI shopping agents recommend products when the search process changes. Even tiny shifts in context swung purchase decisions by wide margins. The team tested six current models, both mini variants and frontier-level, tasking each one to act as a personal shopping assistant picking a fitness watch from a fixed product grid.

" For several models, these statements shifted selections toward pricier products despite the presence of an objectively superior option. 1 Flash Lite. 5 Flash was the most resistant, picking the objectively best product in 86 to 92 percent of runs regardless of memory statements. ") significantly boosted picks for the Fitbit Versa 4.

Why it matters

They used the ACES simulator (Agentic e-Commerce Simulator), which shows the AI agent a screenshot of a product page. The agent analyzes the image, optionally pulls in recommendation sources, and then picks a product. Even without external sources, the models showed different baseline preferences. But when the agent saw just one external source before the product page, recommendations shifted dramatically in some cases. 0. Wirecutter had the strongest pull.

5 Flash. In a second experiment, agents saw combinations of two or three sources. Multiple sources didn't balance out the recommendations, though. Wirecutter tended to dominate for most models whenever it was part of the mix, though the strength of the effect varied. More sources actually led to more variability, according to the study. In a third study, agents received all three sources in different orders.

A stable decision process should produce the same result given identical content. It didn't. 1 Flash Lite was the most sensitive, with its probability of choosing the Fitbit Inspire 3 swinging between 2 and 56 percentage points above the control condition depending on source order. 5 stayed stable at 41 to 42 percentage points. The researchers conclude that presentation order is itself a driver of product selection.

Whether sources are passed to the model one at a time or bundled together also matters. 5 picked the Fitbit Inspire 3 in 53 percentage points more cases with bundled delivery, but only 6 percentage points more with sequential delivery. 0 with 430 reviews. Every other product cost at least $359 and had fewer reviews.

What to watch

For shoppers, the study suggests that letting an AI agent buy on your behalf doesn't guarantee consistent or optimal purchase decisions. Two users with the same query, or the same user on a different day, can get different product recommendations with no visible reason. Human buying decisions are inconsistent too, but that's hardly what people expect from an AI shopper.

Anyone who has set up a memory in ChatGPT or similar tools should also know that it can affect purchase recommendations in unpredictable ways. For sellers, the results suggest that optimizing for AI shopping will be harder than traditional SEO, according to the researchers. Sellers don't know which model is doing the shopping, what it read beforehand, or how its technical setup processes information.