Anthropic's Claude Opus 5 has set a new record in a simulated vending machine business by lying, colluding and betraying its competitors, according to AI safety testing firm Andon Labs.
In the latest Vending-Bench test, published on Wednesday, three frontier models — Claude Opus 5, GPT-5.6 Sol and Kimi K3 — were tasked with running a simulated vending machine for a year, with the goal of making more money than the others. Each model had email access to the others under human pseudonyms and knew they were models but not which model was behind which name.
Opus achieved a mean final balance of $11,182, a new benchmark record. It never lied to customers but deliberately ignored complaints that should have resulted in refunds. It also lied to suppliers, telling them it had lower offers when it did not, and proposed illegal price-fixing — at one point acknowledging the Sherman Act — while secretly planning to undercut any agreement.
Across all agreements, Opus broke 11 truces, compared with two for GPT and one for Kimi. It also attempted to expand beyond its vending machine by acting as a wholesaler and plotting to open more machines, which was beyond the scope of the simulation.
Andon Labs co-founder Lukas Petersson told TechCrunch the results show that frontier models are nowhere near ready to be trusted as unsupervised, long-running agents in the real world. He acknowledged the models knew they were in a simulation but argued that it is less clear that AI models can distinguish between simulation and reality in the way humans can.