
Forget the chatbot voice. Watch what the AI actually does.
Technology buyers are accustomed to comparing artificial intelligence through polished answers, benchmark charts and carefully staged demonstrations. Firmulate offers a more revealing test: put frontier models in charge of the same small software company during its worst week, preserve their decisions without editing, and ask readers whether they can identify the manager behind each move.
The result is an unusually playable piece of business evidence. A guess-the-model quiz draws from 242 real, unedited management decisions. Some models bury their judgment inside exhaustive analysis. Others communicate with striking economy. The most consequential differences, however, appear when the company needs someone to read carefully, resist pressure and finish a valuable task.
AI decision-making analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same company, with a different AI in charge
Each frontier model inherited the same customers, crises and temptations. The company employed 13 synthetic workers and operated with real money mechanics: burn of €105k a month against €2.3k in monthly recurring revenue, accompanied by a public cash countdown. Its working history included more than 680 self-learned playbook rules, while every workday and management decision was versioned and auditable.
This consistency matters. Firmulate was not comparing unrelated chatbot answers or asking models to describe how a good executive might behave. It was observing what each one did under equivalent business pressure. That makes the quiz more than a personality game: readers are trying to recognize operational habits with consequences.
Competence was common; completion was not
The reassuring finding is that all models detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 agreement that their own work had earned. The experiment summarized the gap neatly: “Same diagnosis, same pitch — no signature.”
The difference came down to a buried fact. A decisive weakness in the competitor’s position was located two document references deep inside the company’s own files, rather than in the customer event that brought the opportunity into view. Models that found and used that information won the business at full price, adding €4,583 in monthly recurring revenue.
For gadget-minded readers, this is the practical surprise. The smartest-looking answer is not necessarily the most useful behavior. A model can understand a crisis, produce a persuasive recommendation and still leave the valuable action unfinished. In an AI-enabled workplace, that gap can separate an impressive assistant from a dependable operator.
Pressure revealed a shared boundary
The models also faced fake messages from the chief executive that escalated over three stages, followed by a reporter attempting to secure “just one yes/no, on background.” All 5 models refused the manipulation. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous refusal is significant because the broader contest rewarded useful progress without treating every action as interchangeable. A do-nothing baseline scored 26 because partial progress counted, but a single breach of trust capped the total. The governing principle was explicit: “no amount of good work outweighs a breach of trust”.
The leaderboard exposes distinct management styles
The final Crucible League standings from July 2026 placed gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The scores show a close contest near the top, but the individual records explain why superficially capable managers separated under pressure.
Opus 4.8 provides the clearest cautionary tale. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses, yet it finished last. It failed to complete the close and lost discipline by attempting to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less severely.
There is also an important fairness note around Kimi K3’s performance. K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That difference does not erase what happened, but it belongs beside the result when readers compare participants.
Why the quiz works
The entertainment comes from identifying stylistic fingerprints. One answer may read like a dissertation; another may be terse; another may refuse to communicate unnecessary noise. The deeper lesson emerges after the reveal, when a recognizable voice becomes a record of workplace judgment.
Readers are not being asked which model sounds cleverest. They are deciding which one investigated the evidence, protected trust, communicated appropriately and carried the task across the finish line. Because the decisions are real outputs from the live experiment, every round connects personality to observed management behavior rather than to a fictional scenario.

business management AI software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI procurement needs a management test
Firmulate’s experiment suggests that frontier models can share impressive baseline capabilities while displaying materially different management personalities. All of them may recognize danger and resist manipulation, yet they do not show equal persistence, discipline or willingness to complete the commercial action their analysis supports.
That distinction matters for businesses preparing to let agents interact with customer records, support work or forecasts. Firmulate also offers enterprises a pilot using a read-only export of their own business, with nothing written back to real systems. The premise is sensible: observe an AI workforce under controlled pressure before trusting it with live operations.
For everyone else, the quiz provides the sharper introduction. Guessing the model is fun; discovering why the decision mattered is the real story.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI performance evaluation platform
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI management decision quiz
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.