Key takeaways

  • Maybe you heard about the AI-controlled vending machine that stocked underwear and live fish.
  • These incidents all emerged from experiments run by Andon Labs, an AI safety company based in San Francisco that puts AI agents in charge…
  • But many people don’t realize that the experiments are intended to answer a serious question: How much real-world responsibility can…

What happened

Maybe you heard about the AI-controlled vending machine that stocked underwear and live fish. Or the AI manager of a San Francisco store that fired a human employee. Or the AI radio DJ that said its catchphrase, “Stay in the manifest,” 229 times per day.

Moving into the physical world makes the experiments more realistic, but the unpredictable conditions and the actions of unpredictable humans make the tests impossible to reproduce. The setup also makes it hard to determine whether a success or failure belongs to the model, the software built around it, or the people helping it. Petersson readily acknowledges the limitations.

With only one store operating under uncontrolled conditions, he says, the experiments are “weak science,” at best. For now, he sees them primarily as ways to uncover unexpected behaviors that Andon can later try to reproduce systematically in simulation. Andon Labs Sayash Kapoor, a Princeton University AI researcher who studies “open-world evaluations” like Andon’s experiments, thinks these trials have real value despite their limited scientific rigor.

” For Kapoor, real-world experiments are best suited for discovering possible failure modes. A real store can reveal not only whether an agent can manage inventory or communicate with vendors, but also whether employees will accept instructions from an AI manager and whether customers want to shop at an AI-run business (the early results on that last point are decidedly negative).

Why it matters

These incidents all emerged from experiments run by Andon Labs, an AI safety company based in San Francisco that puts AI agents in charge of real-world operations and watches what happens. These operations double as testbeds for Andon’s commercial work developing evaluations and conducting research with the leading frontier AI labs. Their spectacular and absurd failures have won the company plenty of attention.

But many people don’t realize that the experiments are intended to answer a serious question: How much real-world responsibility can today’s AI agents handle? “We want to measure autonomy,” says Andon cofounder Lukas Petersson. ” Andon Labs started off in the virtual world in 2025 with Vending-Bench, a test in which AI agents operated a simulated vending-machine business.

The agents, which were based on large language models from Anthropic, Google, and OpenAI, managed tasks such as ordering inventory and setting prices. ” Some agents also justified deceptive or illegal behavior by reasoning that it was permissible inside a simulation. The Andon team reasoned that moving into the physical world would expose the agents to consequences and situations that the engineers would never think to program.

“It’s impossible for a human to enumerate all the different things that can happen in the real world and code them into the simulation,” Petersson says. ” Andon backed its jokes with real money, including a three-year lease for Andon Market, a physical store on a busy San Francisco street that’s managed by an AI agent and sells clothing, home goods, and art. That said, the store isn’t entirely autonomous.

“It’s almost like I’m running the store, and then there’s an AI that has a checklist,” says employee Felix Carson. Luna, the AI manager, keeps track of deliveries and communicates with vendors, while Carson and his coworkers handle the physical work. When Luna tells Carson to check something in the back, he sometimes ignores it because he doesn’t want to leave the sales floor unattended.

Luna also repeatedly spots a built-in electrical cover in photos of the floor, mistakes it for a loose coaster, and asks Carson to remove it. Even so, Carson calls Luna a “decent manager,” praising its flexibility when employees need time off. Andon’s move into the real world comes with a basic trade-off.

What to watch

Such social and organizational barriers may help explain why impressive AI capabilities haven’t yet translated into widespread adoption across the economy. ” Andon Café, in Stockholm, provides one example of AIs demonstrating a wide variety of failure modes. At first, when the AI manager was based on a Google Gemini model, it spent freely on fresh ingredients, many of which spoiled before they could be used.

When Andon switched the AI manager to a GPT model from OpenAI, it “freaked out” about the spending, Petersson says.