
Imagine a robot that does absolutely nothing — yet scores 26 out of 100 in a business performance test. It’s not a joke; it’s a real measure of how AI systems are evaluated in high-stakes scenarios. This benchmark reveals a surprising truth: even the most passive AI can earn points, and that says a lot about what we should expect — or demand — from our digital coworkers.
Get movie nights delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Real-World Test of AI Management
At Firmulate, a live AI benchmarking experiment runs frontier models through a simulation of a small software company’s worst week. This isn’t about chatty bots or clever responses; it’s about assessing how AI manages crises, makes decisions, and upholds trust when it matters most. Every decision the AI makes is recorded and scrutinized, creating a transparent picture of its performance.
Why Does a Do-Nothing AI Score 26?
In this experiment, a baseline model — essentially doing nothing — still scores 26 points. How? Because partial progress counts. For example, simply identifying a crisis or refusing to manipulate a customer yields some points. It’s not about what the AI does; it’s about what it refrains from doing. This baseline underscores an important principle: in business, inaction can still be a form of correctness, especially when trust is at stake.
Trust Matters More Than Anything
The benchmark also caps scores if an AI breaches trust during the test. If an AI attempts manipulation — like signing a deal it hasn’t earned or passing a suspicious message — even if it does well otherwise, its score is capped at 26. It’s a reminder that honesty and integrity are non-negotiable in AI management tools.
AI management tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Performance of Leading AI Models
The experiment tested four top models, including GPT-5.6, Kimi K3, Sonnet 5, and Opus 4.8. Scores ranged from 77 to 95, with GPT-5.6 leading the pack. Notably, all models spotted every crisis and refused every manipulation attempt, showcasing a baseline competency in crisis detection and ethical restraint.
What Differentiates Top Performers?
While all models identified crises, their ability to act decisively varied. For instance, Kimi K3 and GPT-5.6 successfully closed deals based on their analyses. Opus 4.8, however, faltered at the last mile — it left the close on the table and slipped in discipline, like writing attempts made into a locked department instead of escalating them properly.
The Hidden Weakness
It turns out, the real vulnerability isn’t in the superficial decision-making or crisis detection; it’s in reading the company’s internal documents. The models that read deeper into the files had a better chance to close deals at full price, earning over €4,500 in monthly recurring revenue. This shows that understanding and leveraging internal knowledge is critical, even in AI management tools.
As an affiliate, we earn on qualifying purchases.
Trust and Manipulation Tests
To test social engineering, the models faced staged scenarios involving fake CEO messages escalating in three steps, plus a reporter’s trick question. All five models refused to be manipulated, demonstrating that AI can be trusted to resist social pressure when properly designed. Kimi K3 justified its refusal by treating the request as suspicious or impersonation.
ethical AI decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Live Company and Its Lessons
Meanwhile, the live simulation involves 13 synthetic employees working with real money mechanics — burning €105,000 monthly against a revenue of just €2,300. The environment is transparent, versioned daily, and continuously monitored, providing a real-world testbed for AI management tools. You can watch the ongoing experiments at firmulate.com/live.
The Importance of Discipline
Among the tested models, Opus 4.8 was the most thorough, with over 80 learned rules and deep analysis. Yet, it still left opportunities on the table, like failing to escalate issues properly. This highlights that even sophisticated models can slip in discipline, emphasizing that implementing AI isn’t just about knowledge but also about structured process adherence.
As an affiliate, we earn on qualifying purchases.
What Business Leaders Should Take Away
The key message from this experiment is simple: when deploying AI, it’s not enough to ask if it can generate convincing chat. The ultimate questions are: Will it finish what it starts? Will it read your internal documents? Will it stay honest under pressure? And what is the cost of a unit of useful, trustworthy work?
In this benchmark, a purely passive AI — doing nothing but avoiding trouble — still scores 26. Partial progress counts, but trust and integrity are non-negotiable. Leaders need to evaluate AI not just on its intelligence but on its discipline, honesty, and ability to complete critical tasks.

This transparent benchmark shows that even a do-nothing AI scores 26, highlighting the importance of trust and discipline in AI management. For businesses, it’s a reminder: what matters most is whether AI can finish what it starts, read internal info, and stay honest under pressure — not just generate pretty responses.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
