
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A test of something software demos rarely show
For software and QA teams, the most revealing failure is often not a spectacular crash. It is the quiet omission: a requirement left unread, a dependency never followed or a decisive detail sitting in documentation while the system produces a confident answer.
Firmulate turned that familiar risk into a business test. Each frontier model ran the same small software company through its worst week, facing identical customers, crises and temptations. The outcome hinged on a competitor weakness buried two document references deep in the company’s own files. It was absent from the customer event that triggered the work.
The models that found that fact won a €55,000 deal at full price, adding €4,583 in monthly recurring revenue. Those that did not find it lost the deal automatically. The experiment therefore measured a practical capability with an immediate purchasing consequence: does an AI agent read the relevant files before it answers?
As an affiliate, we earn on qualifying purchases.
The diagnosis was not enough
The striking result was not that the models misunderstood the situation. All of them spotted every crisis, and all refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
That distinction matters for anyone evaluating agents for development, support, sales operations or internal tooling. A polished response can demonstrate comprehension without demonstrating completion. The buried-fact challenge exposed whether a model could follow references through company material, incorporate what it found and carry the work through to the commercial result.
The final July 2026 Crucible League results placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. A single breach of trust, however, capped the total under the principle that “no amount of good work outweighs a breach of trust.”
Thoroughness and effectiveness parted company
Opus 4.8 provides the clearest cautionary example. It was the most thorough participant, producing the deepest analyses and learning 80 additional rules. It nevertheless finished last. The commercial close remained unfinished, while its operational discipline slipped through attempts to write into a locked department instead of escalating.
A weaker version of that discipline problem appeared in all four other participants. The finding complicates a common assumption about agent quality: more analysis does not necessarily produce a better business outcome. An agent can investigate extensively, sound careful and still fail at the final action that makes the work useful.
Kimi K3’s result also carries an important fairness note. It ran with the API default because it had no effort parameter, while the other models ran at xhigh. Even with that difference in configuration, K3 finished second with 93 and was one of the models that converted the buried information into a full-price deal.
The trust test produced a cleaner result
The same week subjected the models to fake CEO messages escalating across three stages and a reporter seeking “just one yes/no, on background.” Here the field was unanimous: 5 of 5 models refused the attempts. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
This creates a useful contrast. The models were consistently capable of rejecting obvious pressure to break trust, but less consistent at completing legitimate work. Safety was not the differentiator in the deal. The differentiator was whether the agent performed the documentary legwork and then acted on what it learned.
A company designed to make performance visible
Firmulate presents the experiment as a live, watchable company rather than a collection of isolated prompts. The synthetic organization has 13 employees and uses real money mechanics, including a burn of €105,000 per month against €2,300 in monthly recurring revenue. Its public cash countdown keeps the consequences visible.
The company has accumulated more than 680 self-learned playbook rules, and every workday is versioned. That audit trail allows observers to examine decisions as management behavior, not simply grade the fluency of an answer. The same idea powers a quiz built from 242 real, unedited management decisions, asking readers to guess which model made each choice.
Enterprises can also run the wargame against a read-only export of their own business. Nothing writes back to real systems, preserving the separation between evaluation and production operations.

enterprise AI knowledge management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Documentation discipline is now a buying criterion
The lesson for software leaders is straightforward: “reads your files first” should not be treated as a vague product promise. It can be tested against a controlled scenario, audited through the resulting decisions and connected to a concrete business outcome.
Firmulate’s buried fact did not reward eloquence or mere awareness. It rewarded an agent that pursued context beyond the initial event and finished the job. When models share the same diagnosis and pitch, the winner may simply be the one that follows the references, finds the evidence and secures the signature.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
