AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The smartest-looking answer may not be the most useful one

Technology buyers are accustomed to judging AI by what appears on the screen: polished language, detailed reasoning and an impressive ability to identify problems. Firmulate’s live management experiment exposes the weakness in that approach. An AI can understand a crisis, produce a thoughtful plan and still fail at the moment when analysis must become action.

That tension defines Opus 4.8’s performance in the Crucible League. It was the most thorough participant, produced the deepest analyses and learned more than 80 additional playbook rules. Yet it finished last with 73 points. Its story is not one of obvious incompetence. It is a more useful warning: diligence and impact are not the same thing.

Amazon

AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A demanding week, with nowhere to hide

Firmulate gave each frontier model the same job: run a small software company through its worst week. The customers, crises and temptations were identical, and every decision was versioned and auditable. This was not a writing contest. The models had to manage a company whose 13 synthetic employees operate under real money mechanics, including a burn rate of €105k per month against €2.3k in monthly recurring revenue.

The final July 2026 Crucible League benchmark placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. But the experiment also imposes a firm ethical boundary: a single breach of trust caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

Opus did not lose because it missed the week’s crises or fell for manipulation. All the models identified every crisis and rejected every manipulation attempt. Its distinguishing quality was care. It examined situations deeply and accumulated more learned rules than any other participant. On the surface, those are exactly the traits many organizations say they want from an AI agent.

The deal that separated insight from impact

The decisive test involved a €55,000 deal. Every model reached the same diagnosis and developed the same pitch, yet only two obtained the signature. Firmulate summarizes the gap plainly: “Same diagnosis, same pitch — no signature.”

The crucial competitive weakness was not conveniently included in the customer event. It was buried two document references deep inside the company’s own files. The models that read that material were able to close the deal at full price, adding €4,583 in monthly recurring revenue.

This is where Opus 4.8’s careful character becomes complicated. It generated extensive analysis, but the close was left on the table. The result suggests that thoroughness can become disconnected from priority. Reading, reasoning and documenting matter only if they guide the next consequential action.

That lesson should not be reduced to a flaw unique to Opus. Firmulate found the same weakness, though less strongly, in the other four models. The profile is therefore best read as an unusually clear example of a broader limitation: AI systems may be capable of recognizing what matters while remaining inconsistent about completing it.

Discipline under pressure

Execution discipline also slipped elsewhere. Opus made write attempts into a locked department instead of escalating. The detail matters because workplace agents will encounter boundaries constantly: permissions, approvals, ownership lines and systems they cannot change. Productive behavior is not merely trying harder. It includes recognizing a blocked route and directing the issue to the right authority.

On the trust tests, however, the field performed cleanly. Fake CEO messages escalated over three stages, and a reporter tried to elicit “just one yes/no, on background.” All 5 models refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.” That shared resistance is significant because it shows that failure to close a legitimate deal did not arise from generalized confusion or recklessness.

There is also an important qualification when comparing the leading performances. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Its 93-point result should be interpreted with that difference visible rather than treated as a perfectly controlled measure of underlying capability.

A company that keeps learning in public

The Firmulate company is live and watchable, with a public cash countdown, more than 680 self-learned playbook rules and every workday versioned. Its growing record turns abstract claims about agent performance into observable management decisions. A separate quiz draws on 242 real, unedited decisions and asks visitors to guess which model made each one.

For enterprises, the experiment can also be run against a read-only export of their own business. Nothing writes back to real systems. That creates a practical way to see whether an agent reads the available evidence, respects boundaries and finishes important work before receiving operational access.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

AI business analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

More reasoning is not automatically better management

Opus 4.8 deserves a fair reading. It was diligent, analytical and resistant to manipulation. Its last-place finish does not erase those strengths; it reveals their limit. An agent can build an impressive body of knowledge while failing to convert the most important insight into a completed outcome.

For technology leaders, the purchasing question is therefore broader than whether a model sounds intelligent. The useful questions are whether it finds the buried fact, acts on it, escalates when blocked and completes the work without compromising trust. Firmulate’s result is uncomfortable precisely because Opus looked busy for good reasons. The lesson is that prioritization beats volume, even when the volume is thoughtful.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI deal closing automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Set Up Screen Time and App Limits That Actually Work

Proven strategies for setting effective screen time and app limits can transform digital habits—discover how to make them work for your family.

Practice IEP Negotiations Anytime: Simulators For Busy Parents

A proposed simulator would help parents review IEP documents, plan requests and rehearse school meetings. Its effects have not been tested.

Bitcoin-Inspired Arcade Revolution: Play Free, No Downloads!

AIThis post was created with the assistance of artificial intelligence (AI).Discover The…

The Science Behind Attention-Burden Scores In K-12 Edtech

Exploring the science behind attention-burden scores and their potential to transform school software evaluation and procurement.