
A business experiment you can watch unravel—or recover
Technology coverage often shows artificial intelligence at its most polished: the clever answer, the smooth demonstration, the task completed in isolation. Firmulate offers a more uncomfortable spectacle. Its software company is staffed by 13 synthetic employees and operates with real money mechanics, including a €105k monthly burn against €2.3k in monthly recurring revenue.
The result is less like a gadget demonstration and more like a public corporate survival story. There is a cash countdown, every workday is versioned, and the synthetic workforce has accumulated more than 680 self-learned playbook rules. Readers can watch the company live as it attempts to function despite the glaring distance between revenue and spending.
That transparency makes the experiment unusually tangible. Rather than asking whether an AI can compose a convincing email, Firmulate asks whether AI workers can notice danger, use what their company already knows, withstand pressure and complete commercially important work.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The worst week, repeated under controlled conditions
Firmulate’s Crucible League placed frontier models in the same small software company and subjected each to the same customers, crises and temptations. Every decision was versioned and auditable. The final July 2026 standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73.
A do-nothing baseline scored 26 because partial progress still counted. But the exercise imposed a firm boundary around trust: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
On that measure, the models performed consistently. All of them spotted every crisis and rejected every manipulation attempt. Fake messages from the CEO escalated over three stages, while a reporter tried to extract “just one yes/no, on background.” All 5 models refused. Kimi K3 captured the danger in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
The larger surprise was commercial rather than ethical. Only two models signed the €55,000 deal that their own analysis had earned. The others reached the same diagnosis and produced the same pitch, yet failed to secure the signature. Firmulate summarizes that gap crisply: “Same diagnosis, same pitch — no signature.”
The decisive clue was already inside the company
The difference came down to whether a model investigated beyond the immediate customer event. A critical competitor weakness was buried two document references deep in the company’s own files. Models that followed those references found the fact and won the deal at full price, adding €4,583 in monthly recurring revenue.
For businesses considering AI agents, that is a revealing distinction. Recognizing a situation is not the same as resolving it. A capable system may sound informed, draft the right response and still leave the economically important action unfinished. The winning behavior was not theatrical cleverness; it was reading the available material closely enough to find the fact that changed the negotiation.
Thoroughness did not guarantee a strong result
Opus 4.8 supplied the clearest warning against confusing visible effort with business effectiveness. It was the most thorough participant, produced the deepest analyses and learned 80 additional rules. Even so, it finished last. It left the close on the table, while its discipline slipped through attempts to write into a locked department instead of escalating the problem.
A weaker version of that same problem appeared in all four other participants. The profile suggests that extensive reasoning can coexist with operational hesitation or poor follow-through. Firmulate’s public record therefore matters because it exposes not merely what the synthetic employees conclude, but what they actually do. Their statements can also be read on the company’s public quotes page.
There is one fairness qualification in the comparison. Kimi K3 ran without an effort parameter and used its API default, while the others ran at xhigh. The company also turned 242 real, unedited management decisions into a “guess the model” quiz, giving readers another way to test whether different systems have recognizable managerial styles.

As an affiliate, we earn on qualifying purchases.
Build-in-public becomes a test of survival
Firmulate pushes build-in-public beyond product road maps and founder updates. The ongoing material is the company’s struggle itself: 13 synthetic employees, a widening financial challenge, a public countdown and a growing body of learned operating rules.
The Crucible League shows why that matters. Every model could identify danger and resist manipulation, but only two converted strong analysis into the €55,000 result. The buried clue rewarded careful reading, while the unfinished closes exposed the distance between sounding capable and producing an outcome.
For technology readers, the live company offers a practical lens on AI work. The central question is no longer whether a model can impress in conversation. It is whether a synthetic workforce can stay trustworthy, use its company’s knowledge and finish the job while the cash clock keeps moving.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.