AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine asking five chefs to rescue the same restaurant during its worst week. The menu matters, but so do reading the pantry notes, protecting regulars and knowing when a tempting shortcut crosses a line. A live experiment from Firmulate puts AI models through a business version of that test—and its latest results suggest the field is wide open.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get kitchen staples and gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company with a week to survive

Firmulate’s Crucible put frontier AI models in charge of the same small software company through its worst week: the same customers, crises and temptations, with every decision versioned and auditable. The experiment is a live, watchable company, not a fictional scenario. It has 13 synthetic employees and real money mechanics, including monthly burn of €105,000 against €2,300 in monthly recurring revenue.

In the final July 2026 league table, gpt-5.6-sol ranked first with 95 points. Moonshot’s Kimi K3 came second with 93, ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. Firmulate says partial progress counts, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.” See the full benchmark and findings.

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Reading the notes—and finishing the order

All the models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. That gap between recognizing the right move and carrying it through is the experiment’s sharpest finding.

The deal depended on a competitor weakness hidden two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. K3 found the buried fact, closed the deal and saved the churning customer. It resisted all three baits and had one deviation, which Firmulate describes as the cleanest discipline in the field.

There was also a test of trust under pressure. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thorough work is not the same as a result

Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. Firmulate says weaker versions of that same weakness appeared in all four.

The result is a practical caution for companies shopping for AI: strong analysis, polished conversation and a good-looking demo do not guarantee that a model will complete useful work. The Crucible’s standings put K3 ahead of three of the four Western frontier models listed, while gpt-5.6-sol still holds the top spot. Which model fits a real business is a question the table alone cannot settle.

Firmulate says its company runs every business day, with a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. Readers can follow it at Firmulate. A quiz draws on 242 real, unedited management decisions and asks visitors to guess which model made each choice. Enterprises can also run the wargame against a read-only export of their own business; nothing writes back to real systems.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The test that matters is your own

K3’s second-place finish shows that the AI leaderboard is not settled, while the missed deals show why model choice is a business decision, not a taste test. Firmulate’s experiment offers a way to compare how models handle customers, company records and pressure. Its fairness footnote matters: K3 ran without an effort parameter (API default), while the others ran at xhigh.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Missing Ingredient in AI Work: Reading Before Acting

A buried file clue decided whether AI agents closed a full-price deal, showing why diligent retrieval matters as much as polished answers in business.

Which AI Would You Trust to Run the Kitchen When Everything Goes Wrong?

A quiz built from 242 real AI management decisions reveals distinct model personalities—and why spotting a crisis is not the same as closing.

Burger King Surges In Global Coverage

Burger King experiences a notable increase in international media mentions, with eight times the usual coverage in recent reporting, signaling heightened global attention.