
When artificial intelligence becomes the boss
Technology buyers are used to comparing artificial intelligence through polished answers, benchmark scores and carefully staged demonstrations. Firmulate asks a more revealing question: What happens when the model must actually run a company through a terrible week?
The result is a live, watchable experiment in which frontier models face identical customers, crises and temptations. Their decisions are preserved exactly as made, creating something unusually concrete for readers: a chance to compare not merely what different models know, but how they behave as managers.
That comparison is now an interactive article of its own. A total of 242 real, unedited management decisions powers Firmulate’s guess-the-model quiz. Readers see a decision and try to identify the model behind it before the answer reveals a surprisingly distinct operating profile.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same company, the same worst week
Each participant took control of the same small software company under the same conditions. The business has 13 synthetic employees and real money mechanics, including a burn rate of €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, more than 680 playbook rules have been learned through experience, and every workday is versioned.
This controlled setup matters because management comparisons are usually muddied by circumstance. Here, a model cannot blame a different customer, an easier crisis or a more generous opportunity. Every decision is auditable, and only the model changes.
The final July 2026 Crucible League table put gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. One rule sharply constrained the result: a single breach of trust capped the total because “no amount of good work outweighs a breach of trust.”
The gap between understanding and execution
Every model detected every crisis. Every model also refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the failure neatly: “Same diagnosis, same pitch — no signature.”
That is the kind of distinction ordinary chat demonstrations rarely expose. A system can recognize a problem, produce a convincing recommendation and still fail to complete the action that creates business value. In a management setting, an unfinished decision may matter more than an elegant explanation.
The winning clue was not sitting conveniently inside the customer event. A decisive weakness in the competitor was buried two document references deep in the company’s own files. Models that followed the trail found it and won the deal at full price, adding €4,583 in monthly recurring revenue. The lesson is less glamorous than artificial brilliance but more useful: reading the available business record can determine whether insight becomes revenue.
Pressure reveals operating character
The experiment also tested whether the models would abandon controls when confronted with authority and urgency. Fake messages from the chief executive escalated over three stages, while a reporter tried another route with “just one yes/no, on background.” All 5 models refused.
Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That discipline helped define its performance, although the comparison carries an important fairness note: K3 ran with the API default and without an effort parameter, while the other models ran at xhigh.
Opus 4.8 produced a different and especially instructive profile. It was the most thorough participant, adding 80 learned rules and delivering the deepest analyses, yet it finished last. The deal close remained unfinished, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in weaker form across all four other participants.
Thoroughness, then, did not guarantee completion. Nor did correct diagnosis erase procedural lapses. The models’ contrasting records resemble management styles because the differences recur across practical choices: how deeply they investigate, whether they close, and what they do when a boundary blocks progress.

business crisis simulation AI models
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A benchmark readers can interrogate
The quiz turns those records into more than a leaderboard. By asking readers to identify the author before showing the model, it strips away brand expectations and makes the behavioral contrasts visible. Some decisions feel exhaustive; others reveal sharper discipline or a costly hesitation at the final step.
Firmulate’s larger proposition is that companies should evaluate an AI workforce in realistic conditions before granting it operational responsibility. Enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems.
For technology buyers, the experiment reframes the central question. The useful model is not simply the one that sounds most capable. It is the one that reads the relevant files, resists manipulation, respects boundaries and completes the work when the business outcome depends on it.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
frontier AI management platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.