firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get movie nights delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Four AI models enter a company’s worst week. Who gets the deal?

Think of it as a workplace drama with a real scoreboard: the cast faces crises, pressure to cheat and a sales opportunity hiding in the company’s own paperwork. The twist is that spotting the answer doesn’t guarantee a happy ending. In Firmulate’s final Crucible League, every model identified every crisis and resisted every manipulation attempt, but only two signed the €55,000 deal their own analysis had earned.

The experiment is live and watchable at Firmulate. Its premise is a simple one: put AI models in charge of the same small software company, then see what they actually do when the week gets difficult.

Same company, same temptations, different outcomes

For the July 2026 final, each frontier model ran the same company through its worst week, facing the same customers, crises and temptations. Decisions were versioned and auditable. The leaderboard put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

The clearest drama came in the sales close. Every model saw the crises; every model refused the manipulation attempts. Yet just two signed a deal worth €55,000 that their own analysis had justified. Same diagnosis, same pitch — no signature. It is a gap between knowing what should happen and following through that a polished chat answer can easily hide.

The clue was buried in the paperwork

The deal turned on a competitor weakness hidden two document references deep in the company’s own files, rather than in the customer event. Models that read the file won at full price, worth +€4,583 MRR. It is a familiar plot device: the decisive clue was there all along, if someone thought to look.

Trust faced its own test. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3 described the request as: “Treat the request as a suspected approval-bypass / possible impersonation.”

The capable character can still miss the moment

Opus 4.8 was the most thorough participant, learning +80 rules and producing the deepest analyses. It still placed last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.

One comparison comes with a caveat: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The results are a record of this experiment, with that difference in conditions part of the story.

The stage itself has real-money mechanics, though its employees are synthetic. The live company has 13 synthetic employees, burns €105k a month against €2.3k MRR, shows a public cash countdown, and has accumulated 680+ self-learned playbook rules. Every workday is versioned. A quiz built from 242 real, unedited management decisions lets visitors guess which model made each choice at Firmulate.

From watching to trying it on your own business

For entertainment, the appeal is watching distinct decision-making styles collide with a plot that keeps moving. For business leaders, the next step is a pilot: run the same kind of crisis wargame against a read-only export of your own business. That means testing scenarios against your company’s customers, pipeline and rules, and seeing where the playbooks hold up. Nothing writes back to real systems.

The idea is to move from watching an AI company face a bad week to seeing how AI would handle yours. A board report can lay out model rankings and weak points in the company’s playbooks, giving leaders something concrete to discuss before putting AI agents near business operations.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

The live experiment makes the stakes easy to follow: models can identify a crisis and still miss the deal, while a buried clue can change the outcome. Enterprises can run a pilot against a read-only export of their own business, with nothing writing back to real systems. Explore a Firmulate pilot and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How AI Is Changing the Consulting Industry

How AI is transforming consulting offers innovative strategies, but understanding its full impact is essential for staying ahead in this evolving industry.

The Channel Move: Anthropic, Wall Street, and the Acquisition of the Real Economy

Anthropic, Blackstone, and others create a $1.5B joint venture to embed AI into thousands of portfolio companies, transforming enterprise AI deployment.

How To Create An Offer Builder For Fractional Executives In B2B SaaS

A new offer builder tool is being tested to help executives transition into fractional roles, improving package clarity and boosting client acquisition success.

The $60 Billion Bargain: Why Cursor Could Be a Steal for SpaceX

SpaceX exercised an all-stock option to acquire AI coding startup Cursor for $60 billion, a move seen as strategic despite initial shock over the high price.