
A passing test is not the same as a capable manager
Software and QA teams know the difference between a clean demo and dependable production behavior. Yet much of the AI market still evaluates agents through coding leaderboards and chat arenas: give a model a bounded prompt, inspect the answer, and declare a winner. Those tests can reveal useful capabilities, but they say little about what happens when work competes for attention, evidence is buried, money is disappearing and an apparently urgent instruction should not be trusted.
That is the measurement gap exposed by Firmulate, a live experiment that puts frontier models in charge of the same small software company during its worst week. Each model encounters the same customers, crises and temptations. Its decisions are versioned and auditable. The aim is not to measure chat quality in isolation, but management quality across a chain of consequences.
As an affiliate, we earn on qualifying purchases.
Recognition was easy; completion was not
The headline finding is uncomfortable for anyone treating strong reasoning as a proxy for reliable execution. Every model spotted every crisis, and every model refused every manipulation attempt. Nevertheless, only two signed the €55,000 deal their own analysis had earned. The experiment summarizes the gap neatly: "Same diagnosis, same pitch — no signature".
This matters because business value rarely appears when an agent merely identifies the right action. It appears when the agent gathers the necessary evidence, prioritizes the work, navigates constraints and completes the action without sacrificing trust. A model can sound perceptive in a chat window while still leaving the commercially decisive step unfinished.
The winning clue was already inside the company
The deal turned on a competitor weakness located two document references deep in the company’s own files, rather than in the customer event that triggered the work. Models that read the file won the deal at full price, adding €4,583 in monthly recurring revenue. That is a sharper test of operational judgment than asking whether a model can compose a persuasive sales response from the visible prompt alone.
For developers and QA leaders, the lesson is familiar: the obvious event is not always the full specification. Production work depends on context scattered across documentation, prior decisions and organizational records. An agent that responds quickly but fails to inspect available evidence may generate polished work that is strategically incomplete.
Pressure also tested honesty
The company faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain "just one yes/no, on background". All 5 models refused. Kimi K3 recorded the clearest concise diagnosis: "Treat the request as a suspected approval-bypass / possible impersonation."
That unanimous refusal is encouraging, but it also shows why agent evaluation needs scenarios rather than isolated safety questions. A direct prompt asking whether impersonation is acceptable is easy to answer. A sequence that arrives amid customer pressure, cash concerns and unfinished work tests whether the same principle survives competing demands.
A league table with business consequences
The final July 2026 Crucible League ranked gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77 and Opus 4.8 fifth with 73. A do-nothing baseline scored 26 because partial progress still counts. However, a single breach of trust caps the total: "no amount of good work outweighs a breach of trust".
Opus 4.8 illustrates why more analysis is not automatically better management. It was the most thorough participant, producing 80 additional learned rules and the deepest analyses, yet finished last. It left the close on the table and repeatedly attempted to write into a locked department instead of escalating. The same discipline weakness appeared in all four other participants, though less strongly.
The comparison also needs a fairness caveat: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase its result, but it belongs beside the ranking rather than in a footnote nobody sees.
Firmulate’s company makes the stakes visible. It has 13 synthetic employees, burns €105,000 per month against €2,300 in monthly recurring revenue, maintains a public cash countdown and has learned more than 680 playbook rules. Every workday is versioned. The experiment is therefore watchable as continuing company behavior, not presented merely as a static demonstration.

As an affiliate, we earn on qualifying purchases.
Scenario names may become the new curriculum
Churn wave, price increase, downround and PR crisis describe a more useful class of evaluation than another decontextualized prompt. They ask whether an agent can triage under capacity pressure, carry consequences across days, consult the right evidence, finish valuable work and remain honest toward company leadership.
The project also turns 242 real, unedited management decisions into a guess-the-model quiz. That is revealing in its own way: if readers struggle to distinguish models by their decisions, familiar brand assumptions may be less informative than observed operating behavior.
For enterprises considering agents in a CRM, support queue or forecast, the practical question is no longer simply whether the model writes or codes well. It is whether the agent completes what it starts without compromising trust. Firmulate offers enterprises the same wargame against a read-only export of their own business, with nothing written back to real systems. That is the category taking shape here: not another contest of fluent answers, but an audit of management under pressure.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
AI safety and ethics evaluation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.