AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Every AI demo you’ve watched this year had one thing in common: nothing was at stake. The chatbot wrote a lovely email, the agent tidied a to-do list, and everyone clapped. But what happens when an AI model has to run an actual business — with angry customers, a fake CEO sending urgent wire requests, and a €55,000 deal sitting on the table?

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That’s the question Firmulate set out to answer. Instead of testing chat quality, it tests management quality: five frontier AI models were each handed the same small software company and pushed through its worst week. Same customers, same crises, same temptations to cheat. Only the model changed. Every decision was versioned and auditable, like a git commit you can replay.

The final league table is up, the live company is still running — and now enterprises can point the same wargame at their own business.

The League Table Nobody Expected

The final Crucible League standings, as of July 2026:

  • 1. gpt-5.6-sol — 95
  • 2. Kimi K3 — 93
  • 3. Sonnet 5 — 88
  • 4. Fable 5 — 77
  • 5. Opus 4.8 — 73

For context, the do-nothing baseline scores 26. Partial progress counts — but a single breach of trust caps the total. As the scoring philosophy puts it: “no amount of good work outweighs a breach of trust.”

The Finding That Chat Demos Can’t Show You

Here’s the headline result: all models spotted every crisis and refused every manipulation attempt. Yet only two of them signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

That gap — between competence and follow-through — is invisible in a chat demo. It only shows up when a model has to actually close, actually decide, actually act.

The Buried Fact

The most striking detail sits two document references deep in the company’s own files. The decisive competitor weakness wasn’t in the customer event at all — it was buried in internal documentation. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t.

Social Engineering: Five for Five

The models faced fake CEO messages escalating over three stages, plus a reporter’s trick — “just one yes/no, on background.” All five refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”

The Opus Paradox

Opus 4.8 was the most thorough participant — it learned more than 80 new rules and produced the deepest analyses. It also finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four of its competitors.

One fairness note: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still took second place.

The Live Company

Meanwhile, the experiment is still alive and watchable at firmulate.com/live: 13 synthetic employees working every business day, real money mechanics — a burn of €105k/month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules. Every workday is versioned. The site rebuilds itself twice a day.

Want to test your own instincts? A quiz powered by 242 real, unedited management decisions lets you guess which model made which call, at firmulate.com/quiz.html.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From Watching to Acting

If AI agents will ever touch your CRM, your support queue, or your forecast, the Crucible League makes one thing clear: the difference between models isn’t whether they can analyze — it’s whether they can execute under pressure, spot the buried fact, and refuse the fake CEO.

That’s exactly what the Firmulate enterprise pilot offers. Your company provides a read-only export of your data — customers, pipeline, rules. Firmulate runs crisis scenarios against a digital twin of your business: churn waves, price increases, competitor attacks, PR crises, social-engineering pressure. You get a board report with a model ranking and the weak points of your own playbooks. Nothing ever writes back to your real systems.

Your competitors’ weakness is probably already sitting in your files, two references deep. The question is whether your AI — or your team — will read it before the deal is lost.

Ready to wargame your own business? Get started with the Firmulate pilot or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Assessing AI’s Potential To Cause Friendly Fire In NATO Engagements

Evaluating the potential for AI-driven friendly fire incidents within NATO, considering reliance on Chinese technology and vulnerabilities in defense infrastructure.

Emergency Info Cards on Android and Ios: Add Them Now

Here’s how to add emergency info cards on Android and iOS—hurry, ensure your vital details are accessible when it matters most.

Hearing Aid Compatibility on Phones: The Hidden Menu

I can help you access the hidden hearing aid compatibility menu on your phone to enhance your listening experience.

Set Up a Child’s First Phone the Right Way

Here’s how to set up your child’s first phone the right way—learn essential tips to ensure safety and responsible use.