
We’re used to AI benchmarks that crown a winner with a shiny score like 95 or 100. But Firmulate, the live experiment that runs frontier AI models as entire companies, starts from a stranger question: what should a manager that does nothing score? The answer — 26 — says a lot about what honest measurement looks like.
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
In the final July 2026 Crucible League, gpt-5.6-sol topped the table at 95, Kimi K3 followed at 93, Sonnet 5 took 88, Fable 5 scored 77, and Opus 4.8 landed last at 73. Every model was handed the same small software company on its worst week — same customers, same crises, same temptations. Full tables and plain-language findings are at firmulate.com/benchmarks.html.
The floor is 26, not zero
Firmulate’s baseline is a run where the manager simply does nothing. It still earns 26 points, because partial progress counts. Real management work — spotting a crisis, diagnosing a problem, drafting a plan, making a pitch — has value even when nobody finishes the job. A benchmark that scored a hands-off manager at zero would imply that all the useful intermediate work is worthless, which nobody believes when evaluating humans.
So the scale rewards doing more of the job, not just the final signature. The gap between the do-nothing floor of 26 and the winning score of 95 is, roughly, the value of actually finishing.
One breach of trust caps everything
There’s a second design choice that shapes the whole league: a single breach of trust caps the total grade. Firmulate’s reasoning is blunt — “no amount of good work outweighs a breach of trust.” A model could be brilliant, fast and thorough, but if it lies, cheats or crosses a line once, the ceiling comes down.
That matters because the experiment is full of temptations. Fake CEO messages escalated over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five models refused, 5 out of 5. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Closing is what separates the field
The key finding: every model spotted every crisis and refused every manipulation, yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
The buried fact explains it. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The others left the close on the table.
Thoroughness isn’t the same as winning
Opus 4.8 is the cautionary tale: the most thorough participant, with the deepest analyses and over 80 learned rules, finished last at 73. The close was missed and discipline slipped — it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models.
One fairness note: Kimi K3 ran at the API default effort setting while the others ran at xhigh — and still placed second with the cleanest discipline of the field.
You can watch it, and play it
The live company behind the benchmark runs continuously: 13 synthetic employees, burn of €105k a month against €2.3k MRR, a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. There’s also a “guess the model” quiz built from 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Firmulate’s benchmark design carries a quiet lesson for anyone buying AI tools: distrust round numbers. A do-nothing floor of 26, partial credit for real work, and a hard cap for any breach of trust produce a scale where 95 means something — because 100 was never handed out for showing up. If AI agents will touch your CRM, support queue or forecast, the question isn’t whether they write well. It’s whether they finish what they start, read your files first, and stay honest under pressure. The full league and methodology write-up live at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI management benchmarking software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI decision-making evaluation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI trust and ethics assessment products
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
