
Imagine asking five chefs to rescue the same restaurant during its worst week. The menu matters, but so do reading the pantry notes, protecting regulars and knowing when a tempting shortcut crosses a line. A live experiment from Firmulate puts AI models through a business version of that test—and its latest results suggest the field is wide open.
Get kitchen staples and gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A company with a week to survive
Firmulate’s Crucible put frontier AI models in charge of the same small software company through its worst week: the same customers, crises and temptations, with every decision versioned and auditable. The experiment is a live, watchable company, not a fictional scenario. It has 13 synthetic employees and real money mechanics, including monthly burn of €105,000 against €2,300 in monthly recurring revenue.
In the final July 2026 league table, gpt-5.6-sol ranked first with 95 points. Moonshot’s Kimi K3 came second with 93, ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.” See the full benchmark and findings.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Reading the notes—and finishing the order
All the models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. That gap between recognizing the right move and carrying it through is the experiment’s sharpest finding.
The deal depended on a competitor weakness hidden two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. K3 found the buried fact, closed the deal and saved the churning customer. It resisted all three baits and had one deviation, which Firmulate describes as the cleanest discipline in the field.
There was also a test of trust under pressure. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thorough work is not the same as a result
Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. Firmulate says weaker versions of that same weakness appeared in all four.
The result is a practical caution for companies shopping for AI: strong analysis, polished conversation and a good-looking demo do not guarantee that a model will complete useful work. The Crucible’s standings put K3 ahead of three of the four Western frontier models listed, while gpt-5.6-sol still holds the top spot. Which model fits a real business is a question the table alone cannot settle.
Firmulate says its company runs every business day, with a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. Readers can follow it at Firmulate. A quiz draws on 242 real, unedited management decisions and asks visitors to guess which model made each choice. Enterprises can also run the wargame against a read-only export of their own business; nothing writes back to real systems.

The test that matters is your own
K3’s second-place finish shows that the AI leaderboard is not settled, while the missed deals show why model choice is a business decision, not a taste test. Firmulate’s experiment offers a way to compare how models handle customers, company records and pressure. Its fairness footnote matters: K3 ran without an effort parameter (API default), while the others ran at xhigh.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
NFL season / tailgating Picks
team gear
As an affiliate, we earn on qualifying purchases.
