How The Management Test Reveals An AI's True Working Style
AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: How The Management Test Reveals An AI's True Working Style on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

TL;DR

A live management experiment tests AI models on a simulated business crisis, revealing significant differences in diligence, trust, and execution. The results show that analysis alone isn’t enough—effective action is critical. This has important implications for deploying AI in real-world management tasks.

Firmulate’s live management experiment has demonstrated that AI models vary significantly in their ability to not only analyze but also effectively execute business decisions under pressure. The test involved five frontier AI models managing a simulated software company’s worst week, revealing that many models fail to complete critical actions despite accurate diagnoses. This development matters because it challenges assumptions that AI analysis alone suffices for management tasks and highlights the importance of operational discipline. For more on evaluating AI models’ effectiveness, see the original analysis.

The experiment, conducted on firmulate.com, involved five AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8—each tasked with running a small software company through a simulated crisis week. The original analysis of AI decision-making in management can be found here. The models faced identical customer issues, crises, and temptations, with their decisions recorded and auditable. The final leaderboard placed gpt-5.6-sol first with 95 points, while Opus 4.8 scored 73, despite its thorough analysis.

The key finding was that even when models identified the right issues, only two successfully closed critical deals, demonstrating that analysis does not guarantee effective action. Insights into how to assess AI decision-making can be found in this analysis. For example, Opus 4.8 generated extensive insights but repeatedly failed to escalate or complete key operational steps, leading to lower scores. Meanwhile, Kimi K3 refused manipulation attempts, showing strong security instincts, but its overall performance was affected by other operational weaknesses.

One notable result was that models’ ability to recognize risks, such as social engineering attempts, was consistent across participants. All models refused fake CEO requests, indicating they could identify obvious threats. However, differences emerged in their thoroughness, decision execution, and follow-through, which directly impacted their success in closing deals and managing crises.

At a glance
reportWhen: ongoing, with final results from July 2…
The developmentFirmulate’s live experiment pits five AI management models against a simulated business crisis, exposing their strengths and weaknesses in decision-making and execution.

Implications for AI Management and Business Decision-Making

This experiment demonstrates that effective AI management requires more than analysis; models must also reliably execute decisions, escalate issues appropriately, and maintain operational discipline. For enterprises considering AI automation, these findings highlight the importance of testing AI models in realistic, pressure-filled scenarios before deployment. The results suggest that AI’s value in management depends on its ability to act decisively, not just analyze well.

Failing to complete critical operational steps can undermine trust and negatively impact business outcomes, even if the AI’s analysis is sound. As AI models become more integrated into decision-making processes, understanding their working styles and operational discipline will be crucial for safe and effective deployment.

Amazon

AI decision-making management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Testing in Business Management

Traditional AI demonstrations often focus on analysis and language capabilities, with less emphasis on execution and operational discipline. The Firmulate experiment builds on recent efforts to evaluate AI models in real-world management scenarios, using a live, simulated business environment with measurable outcomes. Prior to this, most testing was limited to isolated tasks or hypothetical scenarios, leaving questions about AI’s practical management skills unanswered.

The league table results from July 2026, based on 242 decisions, provide a rare, detailed view of how different models handle complex, pressure-filled management tasks, including risk recognition, trust preservation, and decisive action. This approach emphasizes that AI’s management personality—its diligence, discipline, and follow-through—is as important as its analytical ability.

“Same diagnosis, same pitch — no signature.”

— firmulate.com summary

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unanswered Questions About AI Operational Effectiveness

It is still unclear how these models will perform in real-world, less controlled environments beyond simulations. The experiment focused on a specific business scenario with predefined crises, so results may vary in different contexts or industries. Additionally, the impact of training, customization, and ongoing learning on operational discipline remains to be studied.

Questions also remain about how to best measure and improve AI models’ follow-through capabilities, and whether these findings are consistent across other AI systems and management tasks.

Amazon

AI operational discipline tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI in Business Management Testing

Further research will likely explore deploying these models in actual business settings, with real-time monitoring of their decision-making and execution. Enterprises may adopt similar live testing frameworks to evaluate AI before full integration, reducing risk and ensuring operational reliability. Researchers and developers will also focus on enhancing models’ operational discipline and follow-through skills, aiming to bridge the gap between analysis and action.

Meanwhile, the industry will watch for updates from the ongoing league table and additional experiments that test AI models under different scenarios, to refine understanding of their practical management capabilities.

Amazon

AI crisis management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why is operational discipline important in AI management?

Operational discipline ensures that AI models not only analyze problems but also complete necessary actions, escalate issues properly, and maintain trust—crucial factors for real-world management success.

How does this experiment impact AI deployment in business?

It highlights the need for rigorous testing of AI models in realistic scenarios to assess their follow-through and operational reliability before full deployment.

Can analysis alone guarantee business success with AI?

No, analysis is vital but insufficient; models must also effectively execute decisions and manage operational tasks to succeed in management roles.

What are the limitations of this experiment?

The experiment is based on a simulated scenario, so results may not directly translate to all real-world environments. Further testing in diverse settings is needed.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

When a Content Network Starts Publishing to Itself

A growing trend sees content networks shifting from external distribution to internal publishing, transforming ecosystems and audience control.

Loan covenant calendar for bootstrapped companies

A new workflow for managing loan covenants in small businesses is being tested, focusing on extracting obligations from PDFs to improve compliance.

What Makes a Brand Feel Recession-Proof

Part of making a brand recession-proof involves understanding key strategies that foster lasting resilience—discover what truly makes a brand withstand economic downturns.

Trump’s Latest Gaffes Could Hurt the GOP

Recent comments and missteps by Donald Trump could undermine Republican efforts in upcoming elections, raising concerns among party strategists.