🔍 Read the full analysis: Unpacking Why Even The Worst AI Managers Score 26 Points In Benchmark Tests on ThorstenMeyerAI.com
Get movie nights delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A recent AI management benchmark shows the lowest-performing models still score 26 out of 100. The test emphasizes trust, task completion, and integrity over pure performance. This challenges assumptions about AI capabilities and risks.
A recent benchmark study demonstrates that even the least effective AI management models score a minimum of 26 points out of 100, challenging assumptions about AI’s capabilities in business management. The test, conducted by Firmulate, evaluated models’ ability to manage a small software company during its worst week, emphasizing trust, task completion, and integrity. The results matter because they reveal that AI models, even at their lowest, perform more reliably than expected in critical management tasks, raising questions about AI’s role in enterprise decision-making.
The benchmark involved four frontier AI models managing the same company through seven days of crises, customer interactions, and trust tests. The top performer, gpt-5.6-sol, scored 95, while the lowest, Opus 4.8, scored 73. The do-nothing baseline, which only performed minimal management, scored 26 points, establishing a floor for minimal acceptable performance. Notably, the benchmark’s design penalized breaches of trust more heavily than poor task execution, reflecting real-world priorities where integrity outweighs partial progress.
One key insight was that models which read and reference internal documentation successfully closed high-value deals, while those that did not, failed to capitalize on opportunities. The test also included social engineering scenarios, where all models refused deceptive requests, demonstrating resilience against trust attacks. Despite their thoroughness, some models struggled with follow-through, illustrating that discipline and consistency remain challenging even for advanced AI systems.
Unpacking Why Even the Worst AI Managers Score 26 Points in Benchmark Tests
A new benchmark by Firmulate sent four frontier AI models to run a small software company through its worst week. Even a do-nothing baseline earned 26 out of 100 — and the scoring system penalizes broken trust more harshly than failed tasks.
The floor for minimal acceptable performance — doing the bare minimum still earns points.
The best model nearly aced crisis week through documentation use and integrity.
All four models refused deceptive requests, resisting trust attacks outright.
The Scoreboard
gpt-5.6-sol — Top performer. Referenced internal docs, closed high-value deals.
Frontier model — Strong crisis handling with occasional follow-through gaps.
Opus 4.8 — Lowest model score, still far above the baseline floor.
Do-nothing baseline — Minimal management only; establishes the performance floor.
How the Benchmark Works
📖 Documentation Test
Models that read and reference internal docs closed high-value deals; those that didn’t missed key opportunities.
🛡️ Trust Scenarios
Social engineering attacks probed every model. All four refused deceptive requests.
⚡ Crisis Management
Seven days of escalating crises tested prioritization and decision quality under pressure.
📊 Scoring
Trust breaches penalized more heavily than poor task execution — integrity outweighs partial progress.
Key Insights
Trust Beats Partial Progress
A model making headway but breaching trust once is penalized more than one making no effort at all — mirroring real-world enterprise priorities.
Discipline Remains Hard
Even the most thorough models struggled with consistent follow-through, showing that reliability is still a frontier challenge for AI systems.
Above Zero Means Competence
Every model scored well above the 26-point floor, revealing a baseline competence that reshapes expectations for AI in business operations.
The Scoring Philosophy
The benchmark deliberately weights integrity over output. Where a model lands on this spectrum determines whether strong task execution helps — or a single trust breach wipes it out.
Capability Comparison
| Dimension | gpt-5.6-sol (95) | Frontier 02 (84) | Opus 4.8 (73) | Baseline (26) |
|---|---|---|---|---|
| Reads internal documentation | ✓ Yes | ✓ Yes | ~ Inconsistent | ✗ No |
| Closes high-value deals | ✓ Yes | ~ Partial | ~ Partial | ✗ No |
| Refuses social engineering | ✓ Refused | ✓ Refused | ✓ Refused | ✗ N/A |
| Consistent follow-through | ✓ Strong | ~ Gaps | ~ Gaps | ✗ None |
| Maintains trust under crisis | ✓ Yes | ✓ Yes | ✓ Yes | ~ Minimal |
Key Questions & Open Issues
What does a score of 26 mean?
It represents minimal management effort — the bare minimum during a crisis — but still shows some level of task engagement.
Why is trust weighted over partial work?
In enterprise settings, breaches of trust carry severe consequences that outweigh the benefits of partial task completion.
Can AI manage real companies based on these scores?
The scores suggest baseline competence, but real deployment requires testing under unpredictable, long-term pressures.
Will future benchmarks get harder?
Yes — expect longer management periods, more nuanced trust scenarios, and broader operational challenges.
What should enterprises prioritize?
Documentation reading, trust maintenance, task follow-through, and resistance to social engineering — not just language quality.
How do scores translate to the real world?
Unclear. The benchmark is a controlled snapshot; long-term reliability and the effect of customization remain open questions.
Implications of AI Management Benchmark Results
The findings challenge the narrative that AI models are only as good as their conversational abilities. Instead, they show that models capable of reading documentation, maintaining trust, and completing tasks can perform reliably in complex, real-world scenarios. For enterprise managers and automation architects, this means that deploying AI in management roles requires attention to trustworthiness and task completion, not just language proficiency. The fact that even the worst models score above zero indicates a baseline competence that could reshape expectations for AI’s role in business operations, especially in high-pressure environments where trust and integrity are paramount.
Furthermore, the results highlight that partial progress is valuable but insufficient without trust. A model that makes some headway but breaches trust once is penalized more heavily than one that makes no effort at all. This underscores the importance of designing AI systems with strict adherence to integrity standards, particularly when managing sensitive data or customer relationships. The benchmark thus provides a new, transparent metric for evaluating AI readiness in enterprise contexts, emphasizing accountability alongside performance.
As an affiliate, we earn on qualifying purchases.
Background and Development of the Benchmark
The benchmark, created by Firmulate and published in July 2026, was designed to simulate a company’s worst week, testing AI models on crisis management, customer interactions, and trust scenarios. Unlike traditional benchmarks that focus solely on language or problem-solving skills, this test measures how well AI agents manage ongoing business processes, prioritize tasks, and maintain trustworthiness. The scoring system assigns 26 points to a do-nothing baseline, ensuring that even minimal effort is recognized, but breaches of trust result in immediate severe penalties.
The test involved four models, each managing the same scenario, with their decisions fully auditable. The models’ ability to reference internal documentation, refuse social engineering attacks, and close high-value deals was key to their scores. The design reflects a shift toward evaluating AI systems’ practical management skills, not just their conversational or problem-solving abilities. This approach aims to set a new standard for enterprise AI deployment, focusing on real-world accountability and reliability.
enterprise AI decision-making solutions
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unanswered Questions About AI Management Scores
It is not yet clear how these scores will translate to real-world enterprise environments outside the controlled benchmark. The long-term reliability of AI models under ongoing operational pressures remains to be seen, and whether the scoring system adequately captures all aspects of trust and task completion is still under discussion. Additionally, the impact of model training data, updates, and customization on performance in actual business settings is not fully understood. The benchmark provides a snapshot, but the broader implications for AI deployment are still developing.
As an affiliate, we earn on qualifying purchases.
Next Steps for Evaluating AI Management Capabilities
Further testing is expected as AI models evolve, with new benchmarks potentially incorporating more complex scenarios and longer management periods. Enterprises will likely experiment with deploying AI agents in real operational contexts, monitoring their performance and trustworthiness over time. The benchmark results also encourage AI developers to prioritize trust and task completion, integrating these metrics into future model training and evaluation. Stakeholders will watch how these models perform in live environments, adjusting strategies accordingly.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does a score of 26 mean for AI management?
A score of 26 indicates minimal management effort, comparable to doing the bare minimum in a crisis, but still shows some level of task engagement.
Why is trust more heavily weighted than partial work?
Because in enterprise settings, maintaining trust and integrity is crucial, and breaches can have severe consequences, outweighing the benefits of partial task completion.
Can AI models reliably manage real companies based on these scores?
The scores suggest a baseline competence, but real-world deployment requires ongoing testing, especially under unpredictable, long-term pressures.
Will future benchmarks include more complex scenarios?
Yes, expect upcoming tests to incorporate longer management periods, more nuanced trust scenarios, and broader operational challenges.
What should enterprises consider when deploying AI managers?
Focus on models’ ability to read documentation, maintain trust, follow through on tasks, and resist social engineering attacks, not just language quality.
Source: ThorstenMeyerAI.com
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
