AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A recipe can look flawless on paper and still fall apart in the kitchen. Business leaders face a similar test with AI: a convincing answer is not the same as a sound decision under pressure. Firmulate puts AI models through the rough equivalent of a service gone wrong—using a real, watchable experiment to see how they handle a company’s worst week.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen staples and gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

One company, the same difficult week

In the final July 2026 Crucible League, frontier models ran the same small software company through the same customers, crises and temptations. Every decision was versioned and auditable. The results ranged from 95 for gpt-5.6-sol and 93 for Kimi K3 to 88 for Sonnet 5, 77 for Fable 5 and 73 for Opus 4.8. The do-nothing baseline scored 26. The league’s rule is pointed: partial progress counts, but a single breach of trust caps the total. “No amount of good work outweighs a breach of trust.”

The gap between knowing and doing

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The finding fits in a line: “Same diagnosis, same pitch — no signature.” In a business, as in a busy kitchen, recognizing the problem is only part of the job. Someone still has to carry the decision through.

The decisive clue was not in a customer event. It sat two document references deep in the company’s own files: a competitor weakness. Models that read the file won the deal at full price, worth +€4,583 MRR. The experiment suggests why evaluating an AI on polished answers alone can miss the practical test: does it notice relevant information and act on it?

Pressure, boundaries and a missed close

The models also faced fake CEO messages escalating over three stages, followed by a reporter’s request: “just one yes/no, on background.” All five refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Opus 4.8 offers a more complicated result. It was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four. K3 also ran without an effort parameter, using the API default, while the other models ran at xhigh.

A live test, then a company-specific pilot

Firmulate’s live company has 13 synthetic employees and real money mechanics: burn of €105k a month against €2.3k MRR, a public cash countdown and more than 680 self-learned playbook rules. Every workday is versioned. Readers can watch the experiment at firmulate.com; 242 real, unedited management decisions also power a “guess the model” quiz there.

For enterprises, the next step is a pilot against a read-only export of their own business. The exercise puts crisis scenarios against the company’s customers, pipeline and rules, then provides a board report with model rankings and weak points in its own playbooks. Nothing writes back to real systems. It turns the public experiment’s central question into a company-specific one: how will an AI workforce handle your difficult week?

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your own playbooks to the test

Watch the live experiment, then explore a Firmulate pilot using a read-only export of your business. To discuss a pilot, contact contact@firmulate.com or visit firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

G. L. Mezzetta Inc. Announces Recall Of Mezzetta Brand Golden Greek Peperoncini (Medium Heat) Due To Presence Of Pest Contaminant In The Jar

G. L. Mezzetta Inc. voluntarily recalled 32 oz Golden Greek Peperoncini (Medium Heat), lot 720106, after a pest contaminant report. No illnesses reported.

Norway Became A Global Salmon Behemoth. Now It’s Facing The Consequences

Norway’s rise as a global salmon producer has led to environmental and economic issues, prompting calls for policy changes amid mounting concerns.

Cheesecake Factory Surges In Global Coverage

The Cheesecake Factory has experienced a surge in international media coverage, with 33 mentions in a recent reporting window, highlighting increased global interest.

Will Vermont’s Total Milk Production For 2026 Be Above 2.45 Billion Pounds?

Market activity suggests speculation on Vermont’s 2026 milk output surpassing 2.45 billion pounds, but no official forecast has been confirmed yet.