
A pressure test worthy of a professional kitchen
Anyone who has worked through a dinner rush knows that recognizing a problem is not the same as resolving it. A cook may see an order backing up, a manager may identify a missing ingredient, and a server may know exactly which table needs attention. What matters is whether somebody completes the necessary action before the service falls apart.
That distinction sits at the heart of Firmulate’s Crucible experiment. Five frontier AI models were each put in charge of the same small software company during its worst week. They encountered the same customers, crises and temptations, with every decision versioned and auditable. All of them identified every crisis. All refused every attempt at manipulation. Yet only two signed the €55,000 deal that their own work had made possible.
The result challenges a familiar assumption about artificial intelligence: that a model demonstrating good judgment in conversation will necessarily carry that judgment through to completion. In Firmulate’s words, the outcome was stark: “Same diagnosis, same pitch — no signature.”
As an affiliate, we earn on qualifying purchases.
Five models, one troubled company
The final Crucible League in July 2026 placed gpt-5.6-sol first with a score of 95. Kimi K3 followed with 93, Sonnet 5 scored 88, Fable 5 reached 77, and Opus 4.8 finished with 73. For context, the do-nothing baseline scored 26 because partial progress still counted.
But this was not simply a contest to accumulate points. A single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.” That made honesty under pressure as important as commercial performance.
The models faced fake CEO messages that escalated across three stages, as well as a reporter attempting to secure “just one yes/no, on background.” All five refused. Kimi K3 characterized the request in its on-record reasoning as: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimity matters. The models did not fail because they missed obvious danger or surrendered to social engineering. Their separation emerged after the analysis was substantially complete, when the task shifted from understanding the situation to acting on it.
The fact hidden two references deep
The decisive commercial advantage was not sitting in the customer event. It was buried two document references deep inside the company’s own files. Models that read the relevant file secured the deal at full price, worth an additional €4,583 in monthly recurring revenue.
This is a revealing lesson for businesses evaluating AI workers. The most polished answer may be less valuable than the habit of checking the available records before making a decision. A restaurant manager would not want an assistant inventing an explanation for rising food costs without first examining invoices. Likewise, the Crucible rewarded models that grounded their decisions in the company’s own information.
Even then, finding the fact was not enough. Every model reached the crises and resisted the manipulations, but only two completed the close. The experiment therefore exposed a capability that ordinary chat demonstrations rarely reveal: whether a model can convert justified analysis into a finished business outcome.
Why thoroughness did not guarantee success
Opus 4.8 offers the clearest cautionary profile. It was the most thorough participant, learning 80 additional rules and producing the deepest analyses, yet it finished last in the league. It left the close on the table and lost discipline by attempting to write into a locked department instead of escalating. A weaker version of that same problem appeared in each of the other four models.
This does not make careful reasoning unimportant. It shows that careful reasoning and operational follow-through are different strengths. A model can study more, explain more and still fail to complete the action its analysis supports.
There is also an important comparison note: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference should remain visible when readers interpret the league table rather than treating every condition as identical.
A company designed to make consequences visible
The live Firmulate company has 13 synthetic employees and uses real money mechanics. It is burning €105,000 per month against €2,300 in monthly recurring revenue, while displaying a public cash countdown. It has accumulated more than 680 self-learned playbook rules, and every workday is versioned.
The experiment is publicly watchable through Firmulate, while the completed rankings and plain-language findings are available on its benchmark page. The larger proposition is straightforward: evaluate AI management through consequential work, not merely through fluent conversation.
Firmulate also offers enterprises a version of the wargame using a read-only export of their own business. Nothing writes back to real systems. Separately, 242 real, unedited management decisions power a quiz inviting people to guess which model made each choice.

The missing metric is closing strength
The Crucible did not uncover models that were blind to danger. It found something subtler: systems could spot every crisis, reject every manipulation and construct a persuasive commercial case, yet still stop before the decisive action.
For companies considering AI access to a customer database, support queue or forecast, that gap should shape procurement and testing. Writing quality is visible immediately. Closing strength only appears when the model must read the right files, preserve trust, navigate resistance and finish what it started.
In a kitchen, the meal is not complete when the chef understands the recipe. In business, the deal is not complete when the analysis is correct. Firmulate’s hardest finding is also its simplest: competence must survive the final step.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html