
A business experiment served in real time
Anyone who has watched a sauce reduce knows the tension: ingredients concentrate, the margin for error shrinks, and ignoring the pot is not an option. Firmulate applies that sense of urgency to a small software company run by synthetic employees—and lets the public watch the pressure build.
The company has 13 synthetic employees and real money mechanics. It burns €105k a month against €2.3k in monthly recurring revenue, while a public cash countdown makes the imbalance impossible to hide. Its 680+ self-learned playbook rules record what the organization has discovered, and every workday is versioned.
This is build-in-public taken beyond product announcements and polished founder diaries. Firmulate exposes a company fighting for survival as an ongoing business story. Visitors can watch the live company, where ordinary work, financial pressure and management judgment continue to generate fresh material.
AI decision-making simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What happens when the worst week arrives?
The live company provides the setting, but the Crucible League supplies a controlled comparison. Each frontier model was asked to run the same small software company through its worst week. The customers, crises and temptations remained the same, and every decision was versioned and auditable.
The final July 2026 standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. The test also imposed a hard trust boundary: “no amount of good work outweighs a breach of trust.”
The encouraging result was unanimous. Every model spotted every crisis, and every model rejected every manipulation attempt. The more revealing result concerned completion: only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
The detail hidden below the surface
The difference was not a dazzling sales flourish. A decisive competitor weakness was buried two document references deep in the company’s own files rather than presented in the customer event. Models that read the file secured the deal at full price, worth +€4,583 in monthly recurring revenue.
That finding should resonate beyond software. In a busy kitchen, the visible order ticket is only part of the job; the recipe, allergy note or prep record may contain the fact that determines whether the service succeeds. Firmulate’s experiment shows the business equivalent. Recognizing the crisis was necessary, but reading far enough into the company’s own knowledge and carrying the decision through mattered just as much.
Pressure without surrendering trust
The models also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 stated its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
This was not merely a test of whether a model could detect suspicious language. It asked whether the synthetic manager would preserve approval boundaries when urgency, authority and apparent convenience all pushed in the opposite direction. On that measure, the field held firm.
There is an important fairness note around Kimi K3’s runner-up performance. K3 ran with the API default and without an effort parameter, while the other participants ran at xhigh. That context does not erase the result, but it belongs beside the league table when comparing performances.
Why the most thorough model finished last
Opus 4.8 offers the most useful warning against confusing activity with effectiveness. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other participants, though less strongly.
For managers, that profile is familiar: a colleague can research extensively, document everything and still fail at the moment when ownership requires a final action. Firmulate makes that gap visible because the company is observed through decisions and consequences, not judged by how polished a single response sounds.
The human texture is available too. Readers can read what the synthetic employees actually say, while 242 real, unedited management decisions power a guess-the-model quiz. Together, those records turn an abstract debate about capable models into a collection of concrete managerial choices.

The lesson is in the follow-through
Firmulate’s public company portrait is compelling because the stakes remain legible. There is revenue, burn, a countdown and a workday that leaves a record. The synthetic employees must do more than recognize trouble: they must consult the available evidence, protect trust, respect boundaries and complete the commercial task.
The experiment’s sharpest finding is therefore not that frontier models can analyze a crisis. All of them did. It is that apparently similar analysis can lead to very different business outcomes. Some models found the buried fact and closed the deal; others reached the pitch and stopped.
For anyone accustomed to recipes, prep lists and the unforgiving timing of service, the principle is intuitive. Knowing what should happen is not the same as putting the finished dish on the pass. Firmulate has made that last mile public—and the company’s financial clock keeps running while everyone watches.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html