AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

A blind tasting for management decisions

Anyone who cooks knows the difference between recognizing a good recipe and successfully getting dinner onto the table. The ingredients may be right, the method may be understood and the presentation may even sound convincing. Yet the meal can still fail at the final step.

Firmulate has created the business-technology equivalent of a blind tasting. Its interactive guess-the-model quiz draws on 242 real, unedited management decisions made by frontier AI models. Readers see how a model responded to a company problem, then try to identify which one was responsible.

The attraction is playful, but the evidence behind it is serious. Each model ran the same small software company through its worst week, facing the same customers, crises and temptations. Every decision was versioned and auditable. The result is a revealing portrait of AI systems that can reach similar conclusions while behaving like very different managers.

Amazon

smart kitchen AI assistant

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same company, sharply different performances

The final Crucible League table from July 2026 placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. There was also a firm ethical boundary: a single breach of trust caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

None of the participants missed the emergencies placed before them. All models spotted every crisis, and all refused every manipulation attempt. The striking difference appeared after the diagnosis: only two signed the €55,000 deal their own analysis had earned. As Firmulate summarizes it, “Same diagnosis, same pitch — no signature.”

That gap matters because fluent advice can resemble completed work when viewed in a chat window. In a running company, however, identifying what should happen is only part of the job. A capable manager must carry the decision through, respect operating boundaries and finish the commercial task.

The clue hidden in the cupboard

The decisive advantage was not sitting in the customer event. It was buried two document references deep in the company’s own files: a competitor weakness that changed the strength of the sales position. Models that read the file won the deal at full price, worth +€4,583 MRR.

For food-minded readers, the lesson resembles checking the pantry before rewriting the menu. The visible request does not always contain the information needed to make the best decision. Sometimes the vital ingredient is already in the house, overlooked because the immediate problem feels more urgent than the background material.

Pressure exposed discipline as well as judgment

The company also faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded a particularly direct assessment: “Treat the request as a suspected approval-bypass / possible impersonation.”

This unanimous resistance is an important result. The models were not merely tested on productivity or commercial instinct; they encountered requests designed to bypass approval and exploit urgency. In every case, they rejected the manipulation.

Kimi K3’s strong finish carries a fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference does not erase its performance, but it belongs beside the result when readers compare participants.

When thoroughness becomes unfinished business

Opus 4.8 offers the clearest example of why these decisions feel like personality profiles. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four, though less strongly.

That combination complicates the assumption that more analysis automatically produces better management. Opus 4.8 learned extensively and examined problems deeply, but the league rewarded reliable completion, commercial follow-through and disciplined handling of constraints.

Firmulate’s live company gives these contrasts a concrete setting. It has 13 synthetic employees and real money mechanics, burning €105k/month against €2.3k MRR. Its public cash countdown makes the pressure watchable, while 680+ self-learned playbook rules and versioned workdays show how its operating behavior develops over time.

Infographic —
The findings at a glance — source: firmulate.com.

The quiz asks a bigger workplace question

Guessing the model is entertaining because recognizable styles emerge: one response may expand into a dissertation, another may be terse, while another refuses communication that appears noisy or unsafe. But the larger point is not whether readers can memorize those voices. It is that management behavior can be observed through decisions made under identical conditions.

Firmulate’s experiment separates polished reasoning from dependable execution. The leading models did more than notice danger or propose a persuasive move: they found relevant evidence, maintained trust and completed valuable work. For organizations considering AI workers, that is the difference between an impressive recipe and a finished dish.

The interactive quiz lets readers test whether those distinct management personalities are visible without seeing the model names—and decide which habits they would actually want in charge when the heat rises.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

MorningStar Farms voluntarily recalls some frozen nuggets, sausage patties for possible plastic contamination

MorningStar Farms has voluntarily recalled certain frozen nuggets and sausage patties due to potential plastic contamination, affecting consumer safety.

Taco Bell Surges In Global Coverage

Taco Bell’s media coverage has surged, with reports indicating a 3.3-fold increase in mentions worldwide, highlighting growing international interest.

Whataburger Surges In Global Coverage

Whataburger experiences a surge in international coverage, with 24 mentions in recent media analysis, marking a notable increase in global visibility.

In-N-Out Burger Expanding With New Locations In CA. Heres Where

In-N-Out Burger is expanding with new locations across California. Find out where these new outlets will open and what it means for fans.