
AI can draft a sharp pitch, spot trouble in a customer account and sound confident in a demo. But would it make the decision that actually keeps a business moving? Firmulate put several leading models through the same bad week at a small software company—and found a gap between knowing what to do and doing it.
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company under pressure
Firmulate’s Crucible League gave each model the same customers, crises and temptations. The aim was to see how they handled a company’s worst week, with every decision versioned and auditable. The final results, from July 2026, placed gpt-5.6-sol first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
All the models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The experiment’s sharpest summary was: “Same diagnosis, same pitch — no signature.” It is a revealing distinction for anyone watching AI move from answering questions to handling real business tasks: recognising the right move is not the same as following through.
The clue was buried in the company’s files
The deal turned on a competitor weakness hidden two document references deep in the company’s own files, rather than in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. That detail makes the exercise more than a test of crisis response. The models had to use company knowledge to make a consequential choice.
Firmulate also tested resistance to social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
There are differences behind the leaderboard. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Opus 4.8 was the most thorough participant, adding 80 learned rules and producing the deepest analyses, but finished last. It left the close on the table and tried to write into a locked department instead of escalating. A weaker version of the same discipline problem appeared in all four.
From watching to trying it on your business
The live Firmulate company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k MRR, with a public cash countdown. Its playbook contains more than 680 self-learned rules, and every workday is versioned. Readers can watch the company at firmulate.com. A separate quiz uses 242 real, unedited management decisions to challenge visitors to guess which model made each one.
For business leaders, the next step is a pilot against their own company. Firmulate says an enterprise can provide a read-only data export and run crisis scenarios against a digital twin of its business. The output is a board report with model rankings and the weak points exposed in the company’s playbooks. Nothing writes back to real systems.

Put your own playbooks to the test
Benchmarks can show how models handle a shared challenge. A pilot can reveal what happens when the customers, company knowledge and pressure points are your own. To discuss a Firmulate pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
