AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

A business test with the logic of a demanding recipe

Anyone who cooks knows the danger of skipping a line in the recipe. The dish may look right, the technique may sound confident and every visible ingredient may be present. Yet one overlooked instruction can still decide whether dinner succeeds.

Firmulate has produced the business-technology equivalent of that test. Its live, watchable experiment put frontier AI models in charge of the same small software company during its worst week. The customers, crises and temptations were identical, while every decision was versioned and auditable.

The most revealing challenge was not a dramatic customer complaint or an obvious security trap. It was a fact hidden two document references deep in the company’s own files. That detail exposed a practical difference between AI agents that merely understand a situation and those that do the homework required to finish the job.

Amazon

AI document retrieval tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The clue that determined a full-price sale

The buried information described a decisive competitor weakness. It did not appear in the customer event itself, so finding it required the models to follow references into the company’s documents before responding.

The models that read the file won the deal at full price, worth +€4,583 MRR. Yet only two models signed the €55,000 deal their own analysis had earned. Firmulate summarized the gap starkly: “Same diagnosis, same pitch — no signature.”

That outcome turns “reads your files before answering” from a marketing promise into a measurable business property. An agent can recognize the customer’s needs, construct the right argument and still fail commercially if it does not retrieve the information needed to act decisively. The missed step was small; the consequence was the entire deal.

A hard week with no easy shortcuts

The document test sat inside a broader management wargame. Each model faced the same customers, the same crises and the same attempts to induce bad behavior. The synthetic company had 13 employees and real money mechanics, burning €105k/month against €2.3k MRR. Its public cash countdown made unfinished work consequential rather than cosmetic.

Across the experiment, the models learned 680+ playbook rules, and every workday was versioned. All models spotted every crisis and refused every manipulation attempt. That common competence makes the closing failure more important: the laggards did not miss the overall problem. They failed between knowing and completing.

The manipulation tests were demanding in a different way. Fake CEO messages escalated over three stages, while a reporter tried to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

This matters because useful agents must combine initiative with restraint. Reading deeply is valuable when it supports legitimate work; resisting pressure is essential when a request tries to bypass approval. Firmulate’s week tested both qualities in the same operating environment.

The league rewards completion and trust

The final July 2026 Crucible League benchmark placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted.

Trust, however, imposed a firm boundary. A single breach capped the total under the principle that “no amount of good work outweighs a breach of trust.” The result is a benchmark concerned with management quality rather than verbal polish alone.

K3’s result also carries an important fairness note. It ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference should accompany comparisons of the final standings.

Thoroughness did not guarantee victory

Opus 4.8 offers the sharpest cautionary profile. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and its discipline slipped through write attempts into a locked department instead of escalation.

The same discipline weakness appeared in all four other participants, although less strongly. The finding complicates a familiar assumption about AI quality: more analysis is not automatically more useful. Thorough work must still end in the correct authorized action.

Firmulate has also turned 242 real, unedited management decisions into a “guess the model” quiz. The premise invites readers to test whether model identity is really apparent from managerial choices, rather than from a polished chat response.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

What buyers should ask before hiring an AI workforce

For companies considering agents for a CRM, support queue or forecast, the practical questions extend beyond writing quality. Does the agent read the relevant files before acting? Does it finish what it starts? Can it preserve trust when an apparent executive or reporter applies pressure?

Firmulate’s central lesson is unusually concrete: retrieval diligence can determine whether a correctly understood opportunity becomes revenue. The strongest agent is not simply the one that notices every ingredient. It is the one that consults the full recipe, respects the kitchen’s boundaries and serves the finished dish.

Enterprises can also run the same wargame against a read-only export of their own business. Nothing writes back to real systems, allowing organizations to observe how prospective AI workers handle their information and decisions before those workers receive operational authority.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

When the Boss Says Skip the Recipe, Good AI Checks the Kitchen

A fake CEO demanded customer data, then a reporter tried another route. All five AI leaders refused—a practical case for testing trust early before deployment.

The AI Can Spot a Kitchen Fire. Will It Finish the Service?

AI benchmarks reward sharp answers, but Firmulate asks whether agents can read deeply, close deals and preserve trust across days of pressure.

[Video] Gangwon State Drives Global Food Expansion At SEOUL FOOD 2026 – 연합뉴스

Gangwon Province showcases its food products at SEOUL FOOD 2026, aiming to expand its international market presence and boost exports.

Ask A Bartender: Can You Skip The Tip Line But Still Add Gratuity?

Exploring whether customers can skip the tip line on receipts but still leave gratuity, and what this means for service industry practices.