Unpacking Why Even The Worst AI Managers Score 26 Points In Benchmark Tests
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Unpacking Why Even The Worst AI Managers Score 26 Points In Benchmark Tests on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get movie nights delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

A recent AI management benchmark shows the lowest-performing models still score 26 out of 100. The test emphasizes trust, task completion, and integrity over pure performance. This challenges assumptions about AI capabilities and risks.

A recent benchmark study demonstrates that even the least effective AI management models score a minimum of 26 points out of 100, challenging assumptions about AI’s capabilities in business management. The test, conducted by Firmulate, evaluated models’ ability to manage a small software company during its worst week, emphasizing trust, task completion, and integrity. The results matter because they reveal that AI models, even at their lowest, perform more reliably than expected in critical management tasks, raising questions about AI’s role in enterprise decision-making.

The benchmark involved four frontier AI models managing the same company through seven days of crises, customer interactions, and trust tests. The top performer, gpt-5.6-sol, scored 95, while the lowest, Opus 4.8, scored 73. The do-nothing baseline, which only performed minimal management, scored 26 points, establishing a floor for minimal acceptable performance. Notably, the benchmark’s design penalized breaches of trust more heavily than poor task execution, reflecting real-world priorities where integrity outweighs partial progress.

One key insight was that models which read and reference internal documentation successfully closed high-value deals, while those that did not, failed to capitalize on opportunities. The test also included social engineering scenarios, where all models refused deceptive requests, demonstrating resilience against trust attacks. Despite their thoroughness, some models struggled with follow-through, illustrating that discipline and consistency remain challenging even for advanced AI systems.

At a glance
reportWhen: published July 2026
The developmentA new benchmark testing AI managers’ ability to handle a company’s worst week reveals even the least effective models score 26 points, with significant implications for enterprise AI deployment.
Unpacking Why Even The Worst AI Managers Score 26 Points In Benchmark Tests
AI Management Benchmark · July 2026

Unpacking Why Even the Worst AI Managers Score 26 Points in Benchmark Tests

A new benchmark by Firmulate sent four frontier AI models to run a small software company through its worst week. Even a do-nothing baseline earned 26 out of 100 — and the scoring system penalizes broken trust more harshly than failed tasks.

26 / 100
Do-Nothing Baseline

The floor for minimal acceptable performance — doing the bare minimum still earns points.

95 / 100
Top Score — gpt-5.6-sol

The best model nearly aced crisis week through documentation use and integrity.

0 breaches
Social Engineering Attacks

All four models refused deceptive requests, resisting trust attacks outright.

4
Frontier Models Tested
7
Days of Simulated Crisis
73
Lowest Model Score — Opus 4.8
100%
Refusal of Trust Attacks
3
Core Metrics: Trust · Tasks · Integrity

The Scoreboard

Firmulate Benchmark · Managing a Company’s Worst Week
Model 01
95

gpt-5.6-sol — Top performer. Referenced internal docs, closed high-value deals.

Model 02
84

Frontier model — Strong crisis handling with occasional follow-through gaps.

Model 03
73

Opus 4.8 — Lowest model score, still far above the baseline floor.

Baseline
26

Do-nothing baseline — Minimal management only; establishes the performance floor.

GPT-5.6-SOL
95
FRONTIER MODEL 02
84
OPUS 4.8
73
DO-NOTHING BASELINE
26

How the Benchmark Works

Seven Days · One Company · Fully Auditable Decisions
1

📖 Documentation Test

Models that read and reference internal docs closed high-value deals; those that didn’t missed key opportunities.

2

🛡️ Trust Scenarios

Social engineering attacks probed every model. All four refused deceptive requests.

3

⚡ Crisis Management

Seven days of escalating crises tested prioritization and decision quality under pressure.

4

📊 Scoring

Trust breaches penalized more heavily than poor task execution — integrity outweighs partial progress.

Key Insights

What the Results Reveal About AI in Management
Insight · Integrity

Trust Beats Partial Progress

A model making headway but breaching trust once is penalized more than one making no effort at all — mirroring real-world enterprise priorities.

Insight · Follow-Through

Discipline Remains Hard

Even the most thorough models struggled with consistent follow-through, showing that reliability is still a frontier challenge for AI systems.

Insight · Baseline

Above Zero Means Competence

Every model scored well above the 26-point floor, revealing a baseline competence that reshapes expectations for AI in business operations.

The Scoring Philosophy

How Different Outcomes Weigh on the Final Score

The benchmark deliberately weights integrity over output. Where a model lands on this spectrum determines whether strong task execution helps — or a single trust breach wipes it out.

Do Nothing · 26 pts
Partial Progress · Some Credit
Full Execution + Trust · Max
Trust Breach · Severe Penalty

Capability Comparison

Model Behaviour Across Benchmark Dimensions
Dimension gpt-5.6-sol (95) Frontier 02 (84) Opus 4.8 (73) Baseline (26)
Reads internal documentation✓ Yes✓ Yes~ Inconsistent✗ No
Closes high-value deals✓ Yes~ Partial~ Partial✗ No
Refuses social engineering✓ Refused✓ Refused✓ Refused✗ N/A
Consistent follow-through✓ Strong~ Gaps~ Gaps✗ None
Maintains trust under crisis✓ Yes✓ Yes✓ Yes~ Minimal

Key Questions & Open Issues

What Remains Unanswered

What does a score of 26 mean?

It represents minimal management effort — the bare minimum during a crisis — but still shows some level of task engagement.

Why is trust weighted over partial work?

In enterprise settings, breaches of trust carry severe consequences that outweigh the benefits of partial task completion.

Can AI manage real companies based on these scores?

The scores suggest baseline competence, but real deployment requires testing under unpredictable, long-term pressures.

Will future benchmarks get harder?

Yes — expect longer management periods, more nuanced trust scenarios, and broader operational challenges.

What should enterprises prioritize?

Documentation reading, trust maintenance, task follow-through, and resistance to social engineering — not just language quality.

How do scores translate to the real world?

Unclear. The benchmark is a controlled snapshot; long-term reliability and the effect of customization remain open questions.

Implications of AI Management Benchmark Results

The findings challenge the narrative that AI models are only as good as their conversational abilities. Instead, they show that models capable of reading documentation, maintaining trust, and completing tasks can perform reliably in complex, real-world scenarios. For enterprise managers and automation architects, this means that deploying AI in management roles requires attention to trustworthiness and task completion, not just language proficiency. The fact that even the worst models score above zero indicates a baseline competence that could reshape expectations for AI’s role in business operations, especially in high-pressure environments where trust and integrity are paramount.

Furthermore, the results highlight that partial progress is valuable but insufficient without trust. A model that makes some headway but breaches trust once is penalized more heavily than one that makes no effort at all. This underscores the importance of designing AI systems with strict adherence to integrity standards, particularly when managing sensitive data or customer relationships. The benchmark thus provides a new, transparent metric for evaluating AI readiness in enterprise contexts, emphasizing accountability alongside performance.

Amazon

AI management software tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background and Development of the Benchmark

The benchmark, created by Firmulate and published in July 2026, was designed to simulate a company’s worst week, testing AI models on crisis management, customer interactions, and trust scenarios. Unlike traditional benchmarks that focus solely on language or problem-solving skills, this test measures how well AI agents manage ongoing business processes, prioritize tasks, and maintain trustworthiness. The scoring system assigns 26 points to a do-nothing baseline, ensuring that even minimal effort is recognized, but breaches of trust result in immediate severe penalties.

The test involved four models, each managing the same scenario, with their decisions fully auditable. The models’ ability to reference internal documentation, refuse social engineering attacks, and close high-value deals was key to their scores. The design reflects a shift toward evaluating AI systems’ practical management skills, not just their conversational or problem-solving abilities. This approach aims to set a new standard for enterprise AI deployment, focusing on real-world accountability and reliability.

Amazon

enterprise AI decision-making solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unanswered Questions About AI Management Scores

It is not yet clear how these scores will translate to real-world enterprise environments outside the controlled benchmark. The long-term reliability of AI models under ongoing operational pressures remains to be seen, and whether the scoring system adequately captures all aspects of trust and task completion is still under discussion. Additionally, the impact of model training data, updates, and customization on performance in actual business settings is not fully understood. The benchmark provides a snapshot, but the broader implications for AI deployment are still developing.

Amazon

trustworthy AI management systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Evaluating AI Management Capabilities

Further testing is expected as AI models evolve, with new benchmarks potentially incorporating more complex scenarios and longer management periods. Enterprises will likely experiment with deploying AI agents in real operational contexts, monitoring their performance and trustworthiness over time. The benchmark results also encourage AI developers to prioritize trust and task completion, integrating these metrics into future model training and evaluation. Stakeholders will watch how these models perform in live environments, adjusting strategies accordingly.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does a score of 26 mean for AI management?

A score of 26 indicates minimal management effort, comparable to doing the bare minimum in a crisis, but still shows some level of task engagement.

Why is trust more heavily weighted than partial work?

Because in enterprise settings, maintaining trust and integrity is crucial, and breaches can have severe consequences, outweighing the benefits of partial task completion.

Can AI models reliably manage real companies based on these scores?

The scores suggest a baseline competence, but real-world deployment requires ongoing testing, especially under unpredictable, long-term pressures.

Will future benchmarks include more complex scenarios?

Yes, expect upcoming tests to incorporate longer management periods, more nuanced trust scenarios, and broader operational challenges.

What should enterprises consider when deploying AI managers?

Focus on models’ ability to read documentation, maintain trust, follow through on tasks, and resist social engineering attacks, not just language quality.

Source: ThorstenMeyerAI.com

EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Bitcoin Battles Unfold in Live Warzone Visualization

A new web-based visualization transforms Bitcoin trading data into a cinematic battlefield, offering real-time, immersive insights into market dynamics.

The Twelve Real Complaints About AI Tools in 2026 — A Reddit, Twitter, and GitHub Synthesis

A detailed report on user complaints about AI tools in 2026, highlighting issues like rate limits, context degradation, and reliability concerns from Reddit, Twitter, and GitHub.

Technology Operations Signal Monitor: Explanation Of Everything You Can See In Htop/top On Linux (2019)

A detailed explanation of what the ‘h’ command displays in Linux system monitoring tools like htop and top, for product and engineering leads.

The Ultimate Licensing Hub For Voice Actors’ AI Clones

A new licensing platform for voice actors to control and monetize their AI voice clones has been introduced, aiming to streamline consent, usage, and payments.