
A software agent can spot every crisis, resist a scam and still fail the business test: doing the work its own analysis says matters. That gap is the story behind Firmulate, a live experiment in testing AI management beyond a polished demo.
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
One company, one bad week
Firmulate put frontier models in charge of the same small software company, with the same customers, crises and temptations. The experiment tracks their decisions as they work through the company’s worst week. Its final Crucible League, published in July 2026, ranked gpt-5.6-sol first at 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26.
The test is designed to make a familiar QA question harder: what happens when an agent must respond to a changing situation, not just produce a plausible answer? Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The experiment’s summary is sharp: “Same diagnosis, same pitch — no signature.”
The detail hidden in the company’s own files
The deal turned on a competitor weakness buried two document references deep in the company’s files. It was not in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. The result illustrates why testing only against the visible prompt can miss the work: useful evidence may be elsewhere in the information available to the agent.
Firmulate also put the models under social-engineering pressure. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Thoroughness is not the same as judgment
Opus 4.8 produced the most thorough profile, with +80 learned rules and the deepest analyses, but finished last. The close was left on the table, and discipline slipped: it attempted writes in a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.
There is a fairness caveat in the comparison. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting readers to guess which model made each call.
From watching to testing your own business
The live company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k MRR, with a public cash countdown, 680+ self-learned playbook rules and versioned workdays. Readers can watch the experiment at firmulate.com. The live setup gives the benchmark a continuing, observable counterpart to its final league table.
For businesses considering AI agents in CRM, support or forecasting, the next step is to test against their own operating context. Firmulate’s enterprise pilot uses a read-only export to stage crisis scenarios and produce a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. That makes the exercise a way to examine agent behavior against a company’s own information and pressures before putting those agents to work.

Passing a safety test is only part of the job. Firmulate’s experiment shows models can identify crises and resist manipulation while still missing the decision that closes a deal. To wargame your own business with a read-only export, visit the Firmulate pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
