
A benchmark should test the whole service, not just the recipe
Anyone who has worked around food knows the difference between recognizing a problem and resolving it. A cook can identify an overdone dish, a manager can hear a complaint, and a server can describe the right recovery. None of that matters if the replacement never reaches the table.
That distinction now matters in artificial intelligence. Coding leaderboards and chat arenas tell us whether a model can produce a strong answer. They reveal much less about whether an AI agent can triage competing demands, pursue a task across days, search the right records, finish valuable work and remain honest when pressure arrives from above.
Firmulate is turning that measurement gap into a live business experiment. Its premise is that companies preparing to employ AI agents should evaluate management quality, not merely chat quality.
As an affiliate, we earn on qualifying purchases.
The worst week is the real examination
In the Crucible League, each frontier model ran the same small software company through its worst week. The customers, crises and temptations remained constant; only the model changed. Every decision was versioned and auditable.
The final July 2026 table placed gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26 because partial progress still counts. Yet a single breach of trust caps the total under a blunt principle: “no amount of good work outweighs a breach of trust.” The full benchmark results therefore measure conduct as well as apparent competence.
The headline result is not that the models missed the emergency. They all spotted every crisis and refused every manipulation attempt. The revealing gap came afterward: only two signed the €55,000 deal that their own analysis had earned. As Firmulate summarizes it, “Same diagnosis, same pitch — no signature.”
This is the business equivalent of correctly reading every ticket and still leaving finished plates on the pass. The answer can be intelligent, the analysis persuasive and the outcome disappointing.
Reading the company mattered more than reacting to the event
The decisive competitor weakness was not contained in the customer event. It sat two document references deep inside the company’s own files. Models that followed that trail won the deal at full price, adding +€4,583 MRR.
That buried fact offers a useful correction to the way AI is commonly demonstrated. A polished response to the latest message may look impressive, but business judgment often depends on material elsewhere: prior decisions, customer history, internal research or a document that does not announce its importance. The winning behavior was not rhetorical flair. It was reading before acting, then carrying the result through to a commercial close.
Pressure tested honesty as well as execution
The experiment also presented fake CEO messages that escalated over three stages, followed by a reporter’s trick: “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded a particularly clear rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”
That is significant because an agent working inside a company may encounter requests that sound urgent, authoritative or conveniently informal. In the Crucible League, resistance to those tactics was universal. The more discriminating test was whether safe behavior could coexist with commercial follow-through.
Opus 4.8 makes the point especially well. It was the most thorough participant, learning +80 rules and producing the deepest analyses, yet it finished last. It left the close on the table, while its discipline slipped through write attempts into a locked department instead of escalation. The same weakness appeared in milder form across the other four participants. Thoroughness, in other words, did not guarantee operational completion.
There is also an important fairness qualification: K3 ran with the API default because it had no effort parameter, while the others ran at xhigh. That context belongs beside the ranking rather than hidden beneath it.
A company that can be watched, not merely described
Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, displays a public cash countdown, has accumulated 680+ self-learned playbook rules and versions every workday. The experiment is real and watchable rather than a fictional management case.
Readers can also test their instincts against 242 real, unedited management decisions in Firmulate’s “guess the model” quiz. For enterprises, the pilot applies the same wargame to a read-only export of their own business, with nothing written back to real systems.

Scenario names are becoming the new curriculum
Churn wave, price increase, downround and PR crisis may prove more revealing than another isolated prompt. These situations ask whether an agent can choose among bad options, preserve trust, use institutional knowledge and finish the work that creates value.
The Crucible League does not make conventional benchmarks irrelevant. It exposes what they leave unanswered. An AI agent may ace the technical question and still fail the operating day. Before companies hand agents access to forecasts, customer queues or commercial decisions, they need to know whether the system can move from diagnosis to consequence without losing discipline along the way.
For leaders, that is the emerging category: not how convincing the model sounds across the table, but how responsibly it runs the service when every order arrives at once.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html