AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

A software agent can spot every crisis, resist a scam and still fail the business test: doing the work its own analysis says matters. That gap is the story behind Firmulate, a live experiment in testing AI management beyond a polished demo.

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

One company, one bad week

Firmulate put frontier models in charge of the same small software company, with the same customers, crises and temptations. The experiment tracks their decisions as they work through the company’s worst week. Its final Crucible League, published in July 2026, ranked gpt-5.6-sol first at 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26.

The test is designed to make a familiar QA question harder: what happens when an agent must respond to a changing situation, not just produce a plausible answer? Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The experiment’s summary is sharp: “Same diagnosis, same pitch — no signature.”

The detail hidden in the company’s own files

The deal turned on a competitor weakness buried two document references deep in the company’s files. It was not in the customer event itself. Models that read the file won the deal at full price, worth +€4,583 MRR. The result illustrates why testing only against the visible prompt can miss the work: useful evidence may be elsewhere in the information available to the agent.

Firmulate also put the models under social-engineering pressure. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Thoroughness is not the same as judgment

Opus 4.8 produced the most thorough profile, with +80 learned rules and the deepest analyses, but finished last. The close was left on the table, and discipline slipped: it attempted writes in a locked department instead of escalating. A weaker version of the same weakness appeared in all four models.

There is a fairness caveat in the comparison. Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting readers to guess which model made each call.

From watching to testing your own business

The live company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k MRR, with a public cash countdown, 680+ self-learned playbook rules and versioned workdays. Readers can watch the experiment at firmulate.com. The live setup gives the benchmark a continuing, observable counterpart to its final league table.

For businesses considering AI agents in CRM, support or forecasting, the next step is to test against their own operating context. Firmulate’s enterprise pilot uses a read-only export to stage crisis scenarios and produce a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. That makes the exercise a way to examine agent behavior against a company’s own information and pressures before putting those agents to work.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Passing a safety test is only part of the job. Firmulate’s experiment shows models can identify crises and resist manipulation while still missing the decision that closes a deal. To wargame your own business with a read-only export, visit the Firmulate pilot page or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Neocloud Cartel: How the AI Industry Started Renting Compute From Itself

Exploring how AI companies now rent compute from each other, forming a cartel led by Nvidia, and the implications for industry control and fragility.

Stark Flight Lab Surges In Global Coverage

Stark Flight Lab’s recent surge in coverage, with eight mentions in a short period, indicates rising international interest in its activities and developments.

Device Security In Remote Work: Challenges And Solutions

Exploring how remote teams face device security issues and the emerging solutions to address compliance and security gaps.

Lenovo Surges In Global Coverage

Search interest in Lenovo has surged, with media mentions increasing 9.5 times the baseline, signaling rising global attention. Details on the cause remain unconfirmed.