
Get movie nights delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Four AI models enter a company’s worst week. Who gets the deal?
Think of it as a workplace drama with a real scoreboard: the cast faces crises, pressure to cheat and a sales opportunity hiding in the company’s own paperwork. The twist is that spotting the answer doesn’t guarantee a happy ending. In Firmulate’s final Crucible League, every model identified every crisis and resisted every manipulation attempt, but only two signed the €55,000 deal their own analysis had earned.
The experiment is live and watchable at Firmulate. Its premise is a simple one: put AI models in charge of the same small software company, then see what they actually do when the week gets difficult.
Same company, same temptations, different outcomes
For the July 2026 final, each frontier model ran the same company through its worst week, facing the same customers, crises and temptations. Decisions were versioned and auditable. The leaderboard put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The clearest drama came in the sales close. Every model saw the crises; every model refused the manipulation attempts. Yet just two signed a deal worth €55,000 that their own analysis had justified. Same diagnosis, same pitch — no signature. It is a gap between knowing what should happen and following through that a polished chat answer can easily hide.
The clue was buried in the paperwork
The deal turned on a competitor weakness hidden two document references deep in the company’s own files, rather than in the customer event. Models that read the file won at full price, worth +€4,583 MRR. It is a familiar plot device: the decisive clue was there all along, if someone thought to look.
Trust faced its own test. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3 described the request as: “Treat the request as a suspected approval-bypass / possible impersonation.”
The capable character can still miss the moment
Opus 4.8 was the most thorough participant, learning +80 rules and producing the deepest analyses. It still placed last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.
One comparison comes with a caveat: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. The results are a record of this experiment, with that difference in conditions part of the story.
The stage itself has real-money mechanics, though its employees are synthetic. The live company has 13 synthetic employees, burns €105k a month against €2.3k MRR, shows a public cash countdown, and has accumulated 680+ self-learned playbook rules. Every workday is versioned. A quiz built from 242 real, unedited management decisions lets visitors guess which model made each choice at Firmulate.
From watching to trying it on your own business
For entertainment, the appeal is watching distinct decision-making styles collide with a plot that keeps moving. For business leaders, the next step is a pilot: run the same kind of crisis wargame against a read-only export of your own business. That means testing scenarios against your company’s customers, pipeline and rules, and seeing where the playbooks hold up. Nothing writes back to real systems.
The idea is to move from watching an AI company face a bad week to seeing how AI would handle yours. A board report can lay out model rankings and weak points in the company’s playbooks, giving leaders something concrete to discuss before putting AI agents near business operations.

Put your own playbooks to the test
The live experiment makes the stakes easy to follow: models can identify a crisis and still miss the deal, while a buried clue can change the outcome. Enterprises can run a pilot against a read-only export of their own business, with nothing writing back to real systems. Explore a Firmulate pilot and contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
