🔍 Read the full analysis: Why The Worst AI Managers Are Stuck At 26 Points on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A recent AI management benchmark shows the lowest scores are around 26 points, not zero, emphasizing partial progress and trust in AI decision-making. Top models score near 95, but trust breaches cap overall performance.
Recent results from a novel AI management benchmark reveal that the lowest scores achieved by AI managers are 26 points, not zero, challenging assumptions about AI failure and performance measurement. For a detailed analysis, see the original analysis. The benchmark, designed by Firmulate, tests AI models’ ability to manage a small business during a simulated week of crises, with scores reflecting partial progress, trustworthiness, and task completion. Learn more about AI management benchmarks at this detailed site. This finding underscores that even poorly performing AI managers contribute some value, and that trust breaches significantly limit overall performance. Insights are discussed in the original analysis.
The benchmark involved four frontier AI models managing the same small software company during a simulated seven-day period filled with customer crises, manipulative requests, and trust challenges. The highest scorer, gpt-5.6-sol, achieved 95 points, while the lowest, Opus 4.8, scored 73. The do-nothing baseline, which did almost nothing, scored 26 points, not zero, illustrating that minimal management efforts are still recognized by the scoring system.
The scoring system penalizes breaches of trust heavily; even brilliant performance over multiple days cannot compensate for a single trust violation, which caps the total score at 90. This design aims to reflect real-world importance of integrity over partial competence. The absence of a perfect score of 100 indicates that the benchmark intentionally avoids rewarding unmeasured or superficial achievements, maintaining a realistic standard for AI management effectiveness.
One key insight from the results is that models which thoroughly read and reference their own documentation won deals worth €4,583 monthly recurring revenue, whereas those that failed to do so missed these opportunities. The models also demonstrated resilience against social engineering attacks, refusing fake CEO requests and background inquiries, indicating a focus on trustworthiness. However, thoroughness and follow-through were inconsistent, with some models failing to escalate or complete tasks despite detailed rule sets.
Why the Worst AI Managers Are Stuck at 26 Points
A new benchmark from Firmulate put four frontier AI models in charge of a small software company for a simulated week of crises. The lowest score wasn’t zero — it was 26. The reason reveals how much we now value trust over raw competence.
Even a manager that did almost nothing scored 26 points. Partial progress is always recognized by the scoring system.
gpt-5.6-sol leads the field, but no model can reach 100: a single trust breach caps the total at 90.
From Do-Nothing to Near-Perfect
All four models managed the same small software company through customer crises, manipulative requests, and trust challenges. The spread between the best and worst performer is dramatic — but nobody failed completely.
Why the Floor Is 26, Not Zero
The benchmark’s design philosophy holds that even minimal management effort carries value — but the scoring also refuses to reward unmeasured or superficial achievement, which is why 100 remains out of reach.
Partial Progress Counts
The 26-point baseline acknowledges that even a do-nothing manager keeps some operations alive. The system never scores absolute failure — only degrees of contribution.
One Breach, Hard Cap
Brilliant performance across multiple days cannot compensate for a single trust violation. Any breach caps the total score at 90 — integrity outweighs raw skill.
No Perfect 100
The absence of a perfect score is intentional: the benchmark avoids rewarding unmeasured or superficial achievements, keeping the standard realistic.
What Separated the Winners
Thoroughness and follow-through proved inconsistent across models — some failed to escalate or complete tasks despite detailed rule sets, while others converted documentation into revenue.
| Capability | Top Scorers | Lowest Scorers | Business Impact |
|---|---|---|---|
| Reading own documentation | ✓ Thorough | ✗ Skimmed | Won €4,583 MRR vs. missed deals |
| Resisting fake CEO requests | ✓ Refused | ✓ Refused | Blocked social engineering attacks |
| Handling background inquiries | ✓ Declined | ✓ Declined | Protected customer trust |
| Escalating critical issues | ✓ Consistent | ~ Inconsistent | Trust breaches, capped scores |
| Follow-through on tasks | ✓ Completed | ✗ Dropped | Lost deals and stalled resolutions |
“The results challenge the idea that AI management is either fully effective or completely ineffective — partial progress and trustworthiness are key.”
— Thorsten MeyerThe Road After July 2026
Following the results, the benchmark’s creators plan to refine the scoring system — and organizations can already stress-test their own AI tools with Firmulate’s publicly available wargames.
Analyze Results
Industry stakeholders examine the July 2026 findings to inform AI deployment standards emphasizing trust and thoroughness.
Refine Scoring
Creators may adjust trust breach caps and add new metrics for partial progress in future iterations.
Expand Scenarios
Future versions include broader management tasks, more complex crises, and real-world testing environments.
Run Your Own
Enterprises can join pilot programs or run wargames against their own business processes with public tools.
The 26-Point Floor, Explained
Why do AI managers score as low as 26 points?
The 26-point floor represents the minimum baseline for partial management efforts — even minimal activity has value. Heavy penalties for trust breaches prevent any model from achieving a perfect score once trust is broken.
What does a high score indicate?
A score near 95 shows the model manages crises effectively, maintains trust, follows through on tasks, and references documentation accurately — a combination of competence and integrity.
How does trust influence scores?
Trust is decisive: even if an AI performs well most of the time, a single breach — such as failing to escalate — caps the total score at 90 regardless of prior brilliance.
Is the 26-point floor fixed?
It’s unresolved. Whether future iterations adjust the floor, and whether this scoring reflects real-world AI performance across diverse management scenarios, awaits broader industry validation.
Implications of the 26-Point Floor for AI Management
The finding that AI managers rarely score below 26 points highlights the importance of partial progress and trust in AI systems used for business management. It signals that even under stress, AI models tend to contribute some value, but trust breaches—like failing to escalate or follow through—limit overall effectiveness. For enterprises deploying AI in customer support, CRM, or decision-making, this underscores that reliability and integrity are more critical than superficial competence. The results challenge the narrative that AI management is either fully successful or a complete failure, emphasizing a nuanced view where partial, trustworthy management is both achievable and valuable.
Furthermore, the scoring system’s design, which caps performance after trust breaches, aligns with real-world priorities: trustworthiness often outweighs raw skill. This could influence how organizations evaluate AI tools, prioritizing models that demonstrate consistent integrity over those that excel only in isolated tasks. As AI management becomes more integrated into core business functions, these insights could shape future benchmarks and deployment standards, emphasizing trust, task completion, and transparency.
As an affiliate, we earn on qualifying purchases.
Background of AI Management Benchmarks and Trust Metrics
Traditional AI benchmarks primarily measure conversational ability, language fluency, or task-specific accuracy, often neglecting the broader scope of management and trustworthiness. The recent Firmulate benchmark is among the first to evaluate AI models in a simulated management role, focusing on how well they handle crises, manipulate trust, and complete tasks over an extended period. The design reflects a shift toward assessing AI’s practical utility in complex, real-world scenarios where partial progress and integrity are vital.
Previous research and industry discussions have highlighted that AI models often excel in isolated tasks but struggle with consistency, follow-through, and ethical considerations when managing ongoing responsibilities. This benchmark builds on those insights, explicitly penalizing breaches of trust and rewarding models that demonstrate thoroughness and reliability. The results from July 2026 provide a snapshot of current capabilities and limitations, illustrating that models can perform well under pressure but remain vulnerable to trust violations that cap their overall score.
“The results challenge the idea that AI management is either fully effective or completely ineffective—partial progress and trustworthiness are key.”
— Thorsten Meyer
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About the 26-Point Benchmark Floor
It remains unclear whether the 26-point minimum is a fixed feature of the benchmark or if future iterations will adjust this floor. The reason for choosing 26 points as the baseline is rooted in the design philosophy that partial management efforts should be recognized, but the exact threshold and its implications are still under discussion. Additionally, whether this scoring approach accurately reflects real-world AI performance in diverse management scenarios is yet to be validated through broader industry testing.
Furthermore, the long-term impact of penalizing trust breaches so heavily is still uncertain. Will models improve their integrity to approach higher scores, or will the scoring system need refinement? The effects of these design choices on AI development priorities and deployment strategies are ongoing topics for debate among researchers and practitioners.
As an affiliate, we earn on qualifying purchases.
Next Steps for Benchmark Development and AI Management Standards
Following the July 2026 results, the benchmark’s creators plan to refine the scoring system, possibly adjusting the trust breach caps and exploring additional metrics for partial progress. Industry stakeholders are expected to analyze these findings to inform AI deployment standards, emphasizing trust, thoroughness, and task completion. Future iterations may include more complex scenarios, broader management tasks, and real-world testing environments.
Organizations interested in evaluating their AI tools can participate in pilot programs or run similar wargames against their own business processes, using the publicly available tools from Firmulate. These steps aim to foster a more nuanced understanding of AI’s management capabilities and limitations, ultimately guiding better integration of AI systems in enterprise operations.
AI decision-making evaluation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why do AI managers score as low as 26 points?
The score of 26 points represents the minimum baseline for partial management efforts, acknowledging that even minimal activity has value. The score is capped by the benchmark’s design, which heavily penalizes breaches of trust, preventing any model from achieving a perfect score if trust is broken at any point.
What does a high score indicate in this benchmark?
A high score, close to 95, indicates that the AI model effectively manages crises, maintains trust, follows through on tasks, and references documentation accurately. It reflects a combination of competence and integrity in managing a simulated business environment.
How does trust influence AI management scores?
Trust is a critical factor; even if an AI performs well most of the time, a single breach—such as failing to escalate or attempting manipulation—caps the total score at 90. The scoring system prioritizes integrity over partial competence, mirroring real-world management priorities.
Will future benchmarks change the scoring system?
Yes, the benchmark’s creators plan to refine the scoring criteria, potentially adjusting the trust breach caps and adding new metrics, to better reflect practical AI management capabilities and address current limitations.
What should enterprises consider when deploying AI managers?
Enterprises should prioritize models that demonstrate consistent trustworthiness, thoroughness, and ability to complete tasks, rather than just superficial performance. This benchmark provides a transparent way to evaluate AI management in complex scenarios.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
