firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

In a world obsessed with AI’s shiny capabilities, it’s easy to forget that the most diligent AI isn’t always the most impactful. Imagine four AI models racing to run a virtual company through its worst week—yet only half of them walk away with the big contract. This isn’t science fiction; it’s the real story behind an eye-opening experiment that reveals the surprising limits of pure diligence in AI decision-making.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get movie-night favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Experiment: Putting AI to the Test in a Simulated Business Crisis

Recently, four advanced AI models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—each faced the same grueling scenario: manage a small software company experiencing its worst week. The setup was meticulous, with the same customers, crises, and temptations for dishonest shortcuts. Every decision the models made was tracked, versioned, and auditable, creating a transparent environment to assess how well AI could handle real-world business pressures.

At the end of this trial, the results told a compelling story. All four AI models identified every crisis and refused every manipulation attempt—an impressive feat that underscores their ability to spot trouble and stay honest. Yet, only two managed to close a deal valued at €55,000, earned through their own diagnosis and pitch. The other two, despite their thorough analyses, left the deal on the table, failing to follow through and close the sale.

Amazon

business AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hidden Weakness: The Power of Prioritized Information

Digging deeper, the experiment revealed a crucial insight: the decisive advantage came from reading just two document references deep inside the company’s files—information that was hidden in plain sight. Those models that accessed the file contents successfully won the deal at full price, worth an additional €4,583 in monthly recurring revenue. This highlights a fundamental truth: volume of effort does not automatically translate into impact.

All models demonstrated their ability to recognize crises and resist manipulation, but the difference was in their focus. The models that prioritized reading the company’s own documents—deep, relevant information—were better positioned to make the right call and close the deal. Conversely, the thorough but less strategic models failed to leverage this key insight, leaving money on the table.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Human Angle: Trust and Discipline Under Pressure

In the simulated social engineering test, all models refused fake CEO messages escalating over multiple stages and even a reporter trick asking for a background ‘yes/no’—a testament to their built-in skepticism. Kimi K3, in particular, explained its refusal by labeling such requests as potential impersonation or approval bypasses.

However, the real story isn’t just about refusal. The most detailed model—Opus 4.8—embodying over 80 learned rules and in-depth analysis, still finished last in closing deals. Its weakness lay in discipline lapses: instead of escalating issues, it redirected attempts into locked departments, losing the thread and opportunity. This demonstrates that diligence alone is insufficient if discipline and prioritization aren’t aligned with impactful actions.

Amazon

AI deal-closing automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What This Means for Business and AI

This experiment shatters the myth that volume of effort or exhaustive rules lead to better outcomes. Instead, it underscores a simple but powerful lesson: effective AI in business must prioritize strategic, high-impact information and maintain disciplined focus. AI that ‘reads everything’ may be thorough, but unless it discerns what truly matters, it risks leaving money and trust on the table.

For companies deploying AI—whether in CRM, support, or forecasting—the key takeaway is clear. It’s not about how well an AI can articulate or analyze, but whether it can finish what it starts, stay honest under pressure, and read your most valuable information first. The current leaderboard clearly shows that even the most diligent AI models can stumble if they lack the discipline to prioritize effectively.

Amazon

AI prioritization software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watch the Experiment Live and Decide for Yourself

For those eager to see AI decision-making in action, the live experiment is running at firmulate.com/live. Watch a real software company in the midst of crises, managed by AI models that are constantly learning, versioning, and being tested for their ability to handle real-world business pressures.

In the end, the experiment reveals a fundamental truth: diligence is essential, but it’s not enough. Prioritization and strategic focus are what turn AI from a thorough participant into a game-changer. As AI continues to integrate into critical business functions, understanding its limits and strengths will be key to harnessing its true potential.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Briefro: A Document That Tells The Truth

Briefro introduces an AI-based document tool that keeps data bound to source, runs offline, and guarantees document integrity for regulated industries.

Cameron Whitcomb Surges In Global Coverage

Search and media coverage of Cameron Whitcomb spike significantly, with 21 mentions in the recent window, indicating a surge in international interest.

Contractor onboarding checklist for small construction firms

A new onboarding checklist for small construction firms is being tested to streamline subcontractor onboarding, aiming to reduce delays and administrative gaps.

Comparing Fable, Opus 5.5, Astra, Sol, And Luna: Which AI Model Justifies The Cost?

An in-depth analysis of five leading AI models—Fable, Opus 5.5, Astra, Sol, and Luna—evaluating performance, costs, and suitability for different tasks as of September 2026.