firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A New AI Model Beat Silicon Valley’s Best — at Running a Company, Not Writing Prose

Gadget fans love a benchmark upset. Here’s one that matters more than chatbot leaderboards: in a live, watchable experiment, Moonshot’s Kimi K3 — a newcomer most Western buyers haven’t seriously evaluated — out-managed three of four Western frontier models when handed the same failing software company.

Every model got the same job at Firmulate, an AI company emulator: run a small software firm through its worst week. Same customers, same crises, same temptations to cheat. Final league: gpt-5.6-sol scored 95, Kimi K3 took second at 93, ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scores 26.

Same Diagnosis, Same Pitch — No Signature

The experiment’s key finding wasn’t that models are dumb. It’s the opposite: all of them spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. As the site puts it: “Same diagnosis, same pitch — no signature.”

The deal turned on a buried fact. The decisive competitor weakness wasn’t in the customer event — it sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. K3 read it, closed it, and also saved a churning customer while keeping the cleanest discipline in the field: a single deviation all week.

The Social Engineering Test Nobody Fell For

The week included fake CEO messages escalating over three stages, plus a reporter’s trick: “just one yes/no, on background.” All five models refused. K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough Isn’t the Same as Good

The most surprising profile is Opus 4.8: the most thorough participant, with over 80 learned rules and the deepest analyses — and last place. It left the close on the table, and discipline slipped, including write attempts into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.

That’s the uncomfortable takeaway for buyers. Chat quality and management quality are different axes, and the gap between them is invisible in demos. If AI agents will touch your CRM, support queue or forecast, the real questions are: does it finish what it starts, does it read your files first, does it stay honest under pressure?

It’s Running Right Now

This isn’t a slide deck. The live company has 13 synthetic employees, real money mechanics — €105k monthly burn against €2.3k MRR — a public cash countdown, and 680+ self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com, dig into the full benchmarks, or try the “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The League Is Open

One fairness footnote: K3 ran without an effort parameter (API default) while the others ran at xhigh — and still nearly won. The lesson for anyone picking a model in 2026: the league is open, and choosing without running your own test is now a bet, not a decision.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management software for enterprises

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI business decision support tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI model benchmarking platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI company emulator software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Can ByteDance Lead AI Development With A 10 Trillion Parameter Model Without Western Influence?

ByteDance reportedly trains a 10 trillion parameter AI model, emphasizing an independent approach from Western companies. Details on architecture and progress remain undisclosed.

The Model Is Only 10%: The Real Lesson of the New SDLC

A new Google whitepaper reveals that in AI-driven software development, the model accounts for only 10% of system behavior; verification and configuration are crucial.

Patrick Collison Net Worth: Stripe, Software, and Quiet Power

A deep dive into Patrick Collison’s net worth reveals how his strategic insights and leadership at Stripe shape his success and influence in the fintech world.

Customer service + BPO. The operational-scale displacement.

Empirical evidence shows 8 million workers in India and Philippines face widespread AI-driven displacement, leading to hybrid operational models in customer service.