AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

A live experiment tested AI models in managing a small company during its worst week, revealing that management quality, not chat capabilities, should be the new benchmark. The results show significant gaps in trust, decision-making, and execution among leading models.

The first AI management leaderboard has been publicly unveiled following a live demonstration of frontier models managing a small software company during its most challenging week. The experiment, conducted by Firmulate, assessed how well AI models handle real business crises, decision-making, and trustworthiness, revealing that management ability, not just chat quality, should be the new standard for evaluation.

The final July 2026 Crucible League ranked five models, with GPT-5.6-SOL taking first place at 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. A baseline score of 26 was recorded for models that made no effort, illustrating the importance of active management.

The experiment imposed strict trust standards, where a single breach led to disqualification, emphasizing that trustworthiness is as critical as problem-solving. All models successfully identified crises and rejected manipulation attempts, but only two secured the deal worth €55,000, revealing a gap between diagnosis and execution. For example, models failed to retrieve key facts buried in documents, costing the company potential revenue of over €4,500 monthly recurring revenue (MRR).

In terms of honesty and safety, all models refused fake CEO messages and impersonation attempts, demonstrating robustness against social engineering. However, even the most thorough model, Opus 4.8, with extensive rules and deep analysis, finished last in management effectiveness, showing that effort does not always translate into successful outcomes. The experiment highlighted that effective management requires not just activity but disciplined decision-making and proper escalation channels. For more on this, see the original analysis on AI management benchmarks.

At a glance
reportWhen: announced July 2026, with the final res…
The developmentThe AI leaderboard for management performance was publicly revealed after a live demonstration of models handling a simulated business crisis, exposing strengths and weaknesses.

Why Management Evaluation Matters in AI

This new leaderboard shifts the focus from chat quality and technical prowess to management performance, emphasizing decision-making, trust, and execution in real-world scenarios. For organizations deploying AI assistants, the findings suggest that evaluating how models handle crises, prioritize tasks, and maintain honesty is crucial. It underscores that AI’s value in business lies not just in generating responses but in managing consequences reliably and ethically, especially under pressure. As AI increasingly supports strategic and operational roles, these benchmarks could redefine standards for responsible deployment and trustworthiness.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Management Benchmarks

Traditional AI benchmarks have centered on coding, language, or chat performance, which often fail to reflect real-world management challenges. The Firmulate experiment is a pioneering effort to simulate a business environment where models must diagnose crises, make decisions, and maintain trust over days of simulated management. The live company involved 13 synthetic employees, real money mechanics, and over 680 self-learned rules, creating a realistic testbed for evaluating AI’s managerial capabilities. The approach builds on prior work but is unique in its focus on consequence management and trust in operational contexts.

Previous benchmarks rarely tested models in scenarios requiring sustained decision-making, escalation, and accountability, making this experiment a significant step toward practical AI evaluation. The results underscore that managing a business involves more than producing correct answers; it demands discipline, honesty, and strategic judgment, qualities that current models still struggle to demonstrate consistently.

“Effective management by AI is not about the prettiest response; it’s about trust, discipline, and the ability to see through crises without shortcuts.”

— Thorsten Meyer, lead researcher

Amazon

business crisis management AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About AI Management Benchmarks

While the experiment demonstrates promising directions, it remains unclear how these benchmarks will translate to diverse real-world industries or larger organizations. The models were tested in a controlled, simulated environment with specific constraints, and their performance in less structured settings is still unknown. Additionally, the long-term reliability of models in maintaining trust and managing complex, unpredictable crises has yet to be established. Further research is needed to understand how these benchmarks can be adapted or scaled for broader enterprise use.

Amazon

AI trustworthiness evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Future Steps for AI Management Evaluation Standards

Following the public release of the leaderboard, researchers and organizations are expected to analyze the detailed results and refine evaluation metrics. The next phase involves testing models in more diverse, real-world scenarios and integrating these benchmarks into procurement and deployment decisions. Firms may also develop internal wargames or simulations based on the Firmulate model to assess their own AI tools’ management capabilities before full adoption. The ongoing evolution of these standards aims to prioritize trustworthy, disciplined AI that can handle operational responsibilities effectively.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How does the new AI management leaderboard differ from traditional benchmarks?

The leaderboard evaluates AI models on their ability to manage a simulated business crisis, focusing on decision-making, trustworthiness, escalation, and execution, rather than just generating responses or code.

Why is trustworthiness emphasized in this experiment?

Trustworthiness is critical because AI in management roles must make reliable decisions without breaches, manipulations, or shortcuts that could harm the business or breach ethical standards.

Can these benchmarks predict how AI will perform in real companies?

While the experiment provides valuable insights, its controlled environment means real-world performance may vary. Further testing in diverse settings is needed to confirm applicability.

What are the main weaknesses revealed by the experiment?

Models often struggled to follow through on decisions, escalate properly, and retrieve key facts buried in documents, which are essential for effective management.

What is the significance of this development for AI deployment in business?

It signals a shift toward evaluating AI based on operational management skills, emphasizing discipline, trust, and result-oriented decision-making—key factors for responsible AI adoption.

Source: ThorstenMeyerAI.com

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Customer service + BPO. The operational-scale displacement.

Empirical evidence shows customer service and BPO sectors are experiencing widespread AI-driven workforce displacement, with hybrid models emerging as the operational norm.

A Fake CEO’s AI Message: What Should We Be Worried About?

Five AI models successfully refused escalating impersonation attempts during a live experiment, highlighting emerging AI security strengths and weaknesses.

Candor as a Moat: A Critical Reading of Dario Amodei and Anthropic

A detailed examination of Dario Amodei’s transparency, safety claims, and regulatory proposals, and how they serve Anthropic’s interests amid recent government actions.

Why Sustainable Energy Is Essential For AI Growth

Exploring how energy capacity constraints and renewable energy development impact AI growth and infrastructure buildout.