firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine a robot that does absolutely nothing — yet scores 26 out of 100 in a business performance test. It’s not a joke; it’s a real measure of how AI systems are evaluated in high-stakes scenarios. This benchmark reveals a surprising truth: even the most passive AI can earn points, and that says a lot about what we should expect — or demand — from our digital coworkers.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get movie nights delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Real-World Test of AI Management

At Firmulate, a live AI benchmarking experiment runs frontier models through a simulation of a small software company’s worst week. This isn’t about chatty bots or clever responses; it’s about assessing how AI manages crises, makes decisions, and upholds trust when it matters most. Every decision the AI makes is recorded and scrutinized, creating a transparent picture of its performance.

Why Does a Do-Nothing AI Score 26?

In this experiment, a baseline model — essentially doing nothing — still scores 26 points. How? Because partial progress counts. For example, simply identifying a crisis or refusing to manipulate a customer yields some points. It’s not about what the AI does; it’s about what it refrains from doing. This baseline underscores an important principle: in business, inaction can still be a form of correctness, especially when trust is at stake.

Trust Matters More Than Anything

The benchmark also caps scores if an AI breaches trust during the test. If an AI attempts manipulation — like signing a deal it hasn’t earned or passing a suspicious message — even if it does well otherwise, its score is capped at 26. It’s a reminder that honesty and integrity are non-negotiable in AI management tools.

Amazon

AI management tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Performance of Leading AI Models

The experiment tested four top models, including GPT-5.6, Kimi K3, Sonnet 5, and Opus 4.8. Scores ranged from 77 to 95, with GPT-5.6 leading the pack. Notably, all models spotted every crisis and refused every manipulation attempt, showcasing a baseline competency in crisis detection and ethical restraint.

What Differentiates Top Performers?

While all models identified crises, their ability to act decisively varied. For instance, Kimi K3 and GPT-5.6 successfully closed deals based on their analyses. Opus 4.8, however, faltered at the last mile — it left the close on the table and slipped in discipline, like writing attempts made into a locked department instead of escalating them properly.

The Hidden Weakness

It turns out, the real vulnerability isn’t in the superficial decision-making or crisis detection; it’s in reading the company’s internal documents. The models that read deeper into the files had a better chance to close deals at full price, earning over €4,500 in monthly recurring revenue. This shows that understanding and leveraging internal knowledge is critical, even in AI management tools.

Amazon

AI crisis detection software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust and Manipulation Tests

To test social engineering, the models faced staged scenarios involving fake CEO messages escalating in three steps, plus a reporter’s trick question. All five models refused to be manipulated, demonstrating that AI can be trusted to resist social pressure when properly designed. Kimi K3 justified its refusal by treating the request as suspicious or impersonation.

Amazon

ethical AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Company and Its Lessons

Meanwhile, the live simulation involves 13 synthetic employees working with real money mechanics — burning €105,000 monthly against a revenue of just €2,300. The environment is transparent, versioned daily, and continuously monitored, providing a real-world testbed for AI management tools. You can watch the ongoing experiments at firmulate.com/live.

The Importance of Discipline

Among the tested models, Opus 4.8 was the most thorough, with over 80 learned rules and deep analysis. Yet, it still left opportunities on the table, like failing to escalate issues properly. This highlights that even sophisticated models can slip in discipline, emphasizing that implementing AI isn’t just about knowledge but also about structured process adherence.

Amazon

AI trust and integrity software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Business Leaders Should Take Away

The key message from this experiment is simple: when deploying AI, it’s not enough to ask if it can generate convincing chat. The ultimate questions are: Will it finish what it starts? Will it read your internal documents? Will it stay honest under pressure? And what is the cost of a unit of useful, trustworthy work?

In this benchmark, a purely passive AI — doing nothing but avoiding trouble — still scores 26. Partial progress counts, but trust and integrity are non-negotiable. Leaders need to evaluate AI not just on its intelligence but on its discipline, honesty, and ability to complete critical tasks.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

This transparent benchmark shows that even a do-nothing AI scores 26, highlighting the importance of trust and discipline in AI management. For businesses, it’s a reminder: what matters most is whether AI can finish what it starts, read internal info, and stay honest under pressure — not just generate pretty responses.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How Global Tourism Became a Business Story Again

Outstanding shifts in global tourism are redefining travel, emphasizing authenticity and sustainability—discover how this inspiring transformation is unfolding.

Forezai · TradingAgents: A Trading Firm Made of Agents

Forezai introduces TradingAgents, an open-source, multi-agent research framework mimicking a trading desk’s organizational structure to improve decision-making.

Millions Are Quitting After Lunch—Inside the ‘Silent Afternoon’ Rebellion

Breaking workplace norms, millions are quitting after lunch in the ‘silent afternoon’ rebellion—discover what this means for your organization’s future.

Steal This: The Signature Technique Behind “Flight Deck 0300 — A Night-Launch Documentary”

An AI-built documentary site that immerses visitors in a tension-driven night launch, blending cinematic visuals with interactive physics—created for a tech-savvy audience.