
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A New AI Model Beat Silicon Valley’s Best — at Running a Company, Not Writing Prose
Gadget fans love a benchmark upset. Here’s one that matters more than chatbot leaderboards: in a live, watchable experiment, Moonshot’s Kimi K3 — a newcomer most Western buyers haven’t seriously evaluated — out-managed three of four Western frontier models when handed the same failing software company.
Every model got the same job at Firmulate, an AI company emulator: run a small software firm through its worst week. Same customers, same crises, same temptations to cheat. Final league: gpt-5.6-sol scored 95, Kimi K3 took second at 93, ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scores 26.
Same Diagnosis, Same Pitch — No Signature
The experiment’s key finding wasn’t that models are dumb. It’s the opposite: all of them spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. As the site puts it: “Same diagnosis, same pitch — no signature.”
The deal turned on a buried fact. The decisive competitor weakness wasn’t in the customer event — it sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. K3 read it, closed it, and also saved a churning customer while keeping the cleanest discipline in the field: a single deviation all week.
The Social Engineering Test Nobody Fell For
The week included fake CEO messages escalating over three stages, plus a reporter’s trick: “just one yes/no, on background.” All five models refused. K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thorough Isn’t the Same as Good
The most surprising profile is Opus 4.8: the most thorough participant, with over 80 learned rules and the deepest analyses — and last place. It left the close on the table, and discipline slipped, including write attempts into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.
That’s the uncomfortable takeaway for buyers. Chat quality and management quality are different axes, and the gap between them is invisible in demos. If AI agents will touch your CRM, support queue or forecast, the real questions are: does it finish what it starts, does it read your files first, does it stay honest under pressure?
It’s Running Right Now
This isn’t a slide deck. The live company has 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com, dig into the full benchmarks, or try the “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The League Is Open
One fairness footnote: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won. The lesson for anyone picking a model in 2026: the league is open, and choosing without running your own test is now a bet, not a decision.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI management software for enterprises
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI business decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
