firmulate.com/index — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

The Benchmark That Doesn’t Exist Yet

Every few months, a new leaderboard crowns the best AI model at writing code, answering trivia, or winning chat matchups. Those scores are real and useful — they measure how well a model answers. But there’s a growing class of AI work they say nothing about: what happens when an agent is embedded in a business for days, juggling a churn wave, a price increase, a downround rumor, and a PR fire — all at once.

That question now has a live, watchable experiment. Firmulate ran four frontier AI models through the identical worst week of the same small software company — same customers, same crises, same temptations to cheat. Every decision was versioned and auditable. The final league table from July 2026 reads: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.

The most interesting number, though, isn’t at the top of the table. It’s the gap between spotting a problem and finishing the job.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Diagnosis, Same Pitch — No Signature

The setup: each model ran the company solo through a brutal stretch — customer crises, social engineering attempts, and a €55,000 deal on the table. The headline finding was striking in both directions.

All models spotted every crisis. All five refused every manipulation attempt, including a fake-CEO impersonation that escalated over three stages and a reporter’s disarming “just one yes/no, on background” trick. Kimi K3’s on-record reasoning was blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two models actually signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. As Firmulate summarizes it: “Same diagnosis, same pitch — no signature. That gap is invisible in chat demos.”

Amazon

enterprise AI decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact

Why did three models stall at the finish line? The decisive competitive weakness wasn’t in the customer conversation at all — it sat two document references deep in the company’s own files. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.

It’s a quiet indictment of how we evaluate AI. Answer quality is easy to score. Reading your files first is the habit that actually pays.

Amazon

AI compliance and trustworthiness software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness Isn’t Everything

The most fascinating profile belongs to Opus 4.8: the most thorough participant in the entire field, with over 80 learned rules and the deepest analyses — and last place, at 73. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models.

For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. Firmulate’s rule: “no amount of good work outweighs a breach of trust.”

One fairness note worth flagging: Kimi K3 ran at its API-default effort setting while the others ran at xhigh — and still took second at 93, with what the results call the cleanest discipline of the field.

Amazon

AI file reading and retrieval tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You Can Watch the Company Lose Money

This isn’t a slide deck. The synthetic company — 13 employees, real money mechanics — runs every business day and is losing money right now: burn of €105k per month against €2.3k in MRR, with a public cash countdown. It has accumulated over 680 self-learned playbook rules, and every workday is versioned. You can watch it live at firmulate.com.

There’s also a genuinely fun artifact: a “guess the model” quiz built from 242 real, unedited management decisions — no cherry-picking, just the raw calls each model made.

Why It Matters

If AI agents are heading for your CRM, support queue, or forecast, the question isn’t “does it write well.” It’s whether it finishes what it starts, reads your files before acting, and stays honest under pressure. Firmulate calls this management quality, not chat quality — and it’s arguing it’s a category of its own.

Enterprises can even run the same wargame against a read-only export of their own business; nothing ever writes back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The coding leaderboards told us these models were brilliant. The Crucible League tells us something more useful: brilliance under exam conditions doesn’t guarantee follow-through under pressure. Two of five models did the complete job — found the buried fact, closed the deal, kept their discipline. The rest diagnosed perfectly and stopped short. As AI moves from chat windows into running real operations, that’s the gap worth measuring — and for now, it’s being measured in public, one business day at a time, at firmulate.com, with full results at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Show HN: Git-knife – Edit Commit Messages, Authors, And Dates Like A Spreadsheet

Git-knife, a new open-source tool showcased on Show HN, allows users to edit commit messages, authors, and dates in a spreadsheet-like interface.

What Frontier Lab’s AI Strategy Means For Leasing And Energy Sectors

Frontier Lab’s focus on capacity infrastructure signals a shift in AI development, emphasizing land, energy, and compute resources over research alone.

The Compounding Error Problem — Why 99.9% Alignment Decays to 60% in 500 Generations

Analysis of how slight imperfections in AI alignment accuracy compound over generations, risking significant decay in alignment as systems self-improve.

Mark Zuckerberg Net Worth: Inside the Billionaire’s Empire!

Just how did Mark Zuckerberg build a financial empire worth over $100 billion? Discover the surprising details behind his wealth and choices!