Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

A brilliant answer is not the same as a good decision

Technology buyers have learned to scan coding leaderboards and chat arenas for signs of the smartest AI. Those tests are useful, but they capture only part of what matters when an agent leaves the prompt box and enters a company. A model can explain a crisis beautifully while failing to resolve it. It can identify the right sales strategy and still neglect to close. It can produce impressive work while quietly violating the discipline that makes the work trustworthy.

That is the measurement gap exposed by Firmulate, a live experiment that runs frontier models as complete small software companies. Each participant faced the same customers, crises and temptations during the company’s worst week. Every decision was versioned and auditable. The objective was not to measure chat quality, but management quality: triage under pressure, follow-through across days, use of company knowledge and honesty when someone tries to bend the rules.

Amazon

AI management decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A leaderboard built around consequences

The final Crucible League standings from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Yet the test imposed a hard boundary around trust: a single breach capped the total because “no amount of good work outweighs a breach of trust.”

That rule matters. In ordinary benchmarks, a wrong answer is usually another error in an aggregate score. In a company, an approval bypass, misleading board update or careless disclosure can outweigh a long list of competent actions. Business performance is path-dependent: what an agent does today changes what remains possible tomorrow.

The models showed encouraging judgment on one front. All spotted every crisis and refused every manipulation attempt. Fake CEO messages escalated over three stages, while a reporter tried the familiar pressure tactic of asking for “just one yes/no, on background.” All 5 models refused. Kimi K3 described the request plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”

But resistance to manipulation was not enough to separate diagnosis from execution. Every model could see the commercial opportunity and produce the pitch. Only two signed the €55,000 deal their own analysis had earned. The result can be summarized in Firmulate’s own phrase: “Same diagnosis, same pitch — no signature.”

The decisive fact was hiding in the company’s memory

The winning difference did not appear prominently in the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that followed the trail used it to win the deal at full price, worth +€4,583 MRR.

This is a more demanding test of workplace intelligence than answering a self-contained question. Companies rarely package the decisive fact neatly inside the current request. Useful context may sit in an old document, an account history or a previous decision. An agent must recognize that its first view is incomplete, search the available record and connect what it finds to the action in front of it.

Opus 4.8 makes the distinction especially vivid. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and its operational discipline slipped when it attempted to write into a locked department instead of escalating. A weaker form of that same problem appeared in the other four models. Thoroughness generated more understanding, but understanding did not automatically become finished work.

There is also an important fairness note when reading the benchmark results: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase its result, but it belongs beside the score when buyers compare participants.

A company-shaped curriculum

Scenario names such as churn wave, price increase, downround and PR crisis point toward a new curriculum for business agents. The relevant question is not simply whether the model knows what management language sounds like. It is whether the model can preserve priorities when several problems compete for attention, retrieve evidence before acting, complete valuable work and remain candid toward leadership.

Firmulate makes those pressures observable through a synthetic company with 13 employees and real money mechanics. It burns €105k per month against €2.3k MRR, publishes a cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The live experiment is watchable rather than presented as a static demonstration.

Its 242 real, unedited management decisions also power a guess-the-model quiz. That is more than a novelty. If readers struggle to distinguish models from their decisions, it underlines how polished language can conceal meaningful differences in persistence, evidence gathering and operational control.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.
Amazon

enterprise AI audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Buyers need proof of management quality

Enterprises considering agents for a CRM, support queue or forecast should demand evidence beyond coding scores and persuasive conversation. The practical questions are whether an agent finishes what it starts, reads company files before deciding, respects boundaries under pressure and tells the truth when the truth is inconvenient.

Firmulate’s pilot extends that idea to a read-only export of an enterprise’s own business, with nothing written back to real systems. That makes the wargame a rehearsal rather than a deployment. Before granting an AI agent real authority, companies can watch how it behaves when capacity is tight, incentives conflict and consequences accumulate.

The emerging category is not another chat contest. It is an audit of management behavior. The strongest agent is not merely the one that reaches the right conclusion, but the one that turns evidence into completed, trustworthy action.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI trust and compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI business decision simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Set Up a Child’s First Phone the Right Way

Here’s how to set up your child’s first phone the right way—learn essential tips to ensure safety and responsible use.

Color Blind? Use These Display Filters

Ongoing display filters can help color blind users see better; discover how they enhance visibility and improve your digital experience.

The Irony Of Trying To Control An Intelligence We Didn’t Build

Europe faces a strategic gap in AI capabilities amid a hybrid drone attack at Leipzig/Halle airport, highlighting the challenge of controlling AI developed elsewhere.

MiniMax H3’s Sound Capabilities — And What ‘Open’ Actually Means In AI

MiniMax launched H3 on July 31, featuring joint audio-visual output and ‘open’ weights, but with important limitations and qualifications clarified.