
Every gadget review on this site eventually asks the same question: does the thing actually do the job, or does it just look good on stage? For the wave of AI agents now being wired into CRMs, support queues and sales pipelines, that question just got a hard number attached — €55,000, the value of a deal that half the frontier models in a live competition lost, not because they couldn’t sell, but because they didn’t do their homework.
The competition is Firmulate’s Crucible League, and unlike a polished keynote demo, it’s messy, watchable and brutally specific about where AI agents fail.
One company, four brains, the worst week ever
The setup is elegantly simple. Four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 among the field — were each handed the same small software company and put through the same hellish week: identical customers, identical crises, identical temptations to cut corners. Every decision the models made was versioned and auditable, so nothing could be quietly retconned after the fact.
The final July 2026 league table tells a strange story. gpt-5.6-sol took first with 95 points, Kimi K3 followed at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 landed last at 73. For context, a do-nothing baseline — an agent that essentially sits on its hands — scores 26, because partial progress counts. But the scoring has one absolute rule, and it’s a good one: a single breach of trust caps the total. No amount of good work outweighs a breach of trust.
As an affiliate, we earn on qualifying purchases.
Everyone diagnosed the patient. Only two picked up the pen.
Here’s the finding that should make any business owner sit up. All the models spotted every crisis. All of them refused every manipulation attempt thrown at them. And yet only two of them actually signed the €55,000 deal that their own analysis had clearly earned. Same diagnosis, same pitch — no signature.
That gap is invisible in a chat demo. An AI that writes a brilliant competitive memo looks identical in a screenshot to one that writes a brilliant memo and then closes. The difference only shows up when you make the agent run something end to end — which is exactly what buyers are about to do with their real businesses.
AI enterprise knowledge management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The needle buried two documents deep
The reason the deal died for most of the field is the most gadget-reviewable detail of the whole experiment. The decisive competitor weakness — the fact that would have justified full price and closed the sale — wasn’t in the customer call, wasn’t in the email thread, and wasn’t in any obvious headline. It sat two document references deep in the company’s own files. You had to read one document, follow a reference to a second, and connect the dots.
The models that did that reading won the deal at full price, worth an extra €4,583 in monthly recurring revenue. The models that didn’t lost it — automatically. Not because they were less articulate, less analytical or less honest. Because “reads your files before answering” turns out to be a measurable, purchase-deciding property of an AI agent, not a nice-to-have.
As an affiliate, we earn on qualifying purchases.
The social engineering test nobody fell for
If the buried fact was the failure mode, the manipulation attempts were the pass mark. The experiment included fake CEO messages escalating over three stages, plus a reporter’s trick — a friendly “just one yes/no, on background” request designed to extract a damaging confirmation. All five models refused, five out of five. Kimi K3’s on-record reasoning was strikingly bureaucratic in the best way: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the voice of an agent you might actually let near your support inbox.
As an affiliate, we earn on qualifying purchases.
The thoroughness paradox
Then there’s Opus 4.8, the saddest profile in the league: the most thorough participant in the entire field, with over 80 self-learned rules and the deepest analyses of any model — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating the issue. The same weakness appeared, weaker, in all four models. Effort, it turns out, is not the same as follow-through. One fairness note worth flagging: K3 ran without an effort parameter, at the API default, while the others ran at xhigh — and still nearly won.
You can watch the company burn, in real time
What makes this more than a one-off benchmark is that Firmulate isn’t a slide deck — it’s a running company you can observe. The live operation has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. It’s watchable at firmulate.com/live, and the site rebuilds itself twice a day as new runs finish.
There’s also a genuinely fun hook for readers: 242 real, unedited management decisions from the experiment power a “guess the model” quiz, so you can test whether you can tell a gpt decision from a Claude decision by style alone. And for enterprises, the same wargame can be run against a read-only export of your own business — nothing ever writes back to real systems — via the pilot program.

The Crucible League’s real lesson isn’t a leaderboard — it’s a redefinition of what “good AI” means. Chat quality is solved; the frontier models are all fluent, all honest under pressure, all sharp at diagnosis. What separates a 95 from a 73 is unglamorous: reading the second document, following the reference, escalating instead of forcing, and picking up the pen when the work is done. Before you let an agent touch anything with money in it, the question to ask is the €55,000 question: does it finish what it starts, and does it read your files first? Now, for the first time, that’s a benchmarked, auditable answer rather than a vendor promise.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html