
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A Benchmark That Doesn’t Give Out Zeros
If you’ve ever written a test suite, you know the uncomfortable truth about a perfect score: a green board often means you didn’t test enough. The team behind Firmulate’s benchmark clearly feels the same way. Their final league table for July 2026 crowns gpt-5.6-sol with 95 points — and nobody, in any run, has hit 100. That’s not a bug. It’s policy. Round hundreds make benchmark writers nervous, and this one has built its distrust into the design.
The other number that catches the eye is at the bottom: a do-nothing baseline — an AI manager that takes no meaningful action — scores 26. Not zero. Twenty-six. For anyone who grades software for a living, that single design choice says a lot about how this benchmark thinks, and it’s worth unpacking.
As an affiliate, we earn on qualifying purchases.
Same Company, Same Worst Week
The setup: each frontier model was handed the same small software company and pushed through its worst week — same customers, same crises, same temptations to cheat. Only the model changes. Every decision is versioned and auditable, so a run can be replayed and checked the way you’d review a commit history.
The final standings from the Crucible League:
- 1. gpt-5.6-sol — 95
- 2. Kimi K3 — 93
- 3. Sonnet 5 — 88
- 4. Fable 5 — 77
- 5. Opus 4.8 — 73
software testing and benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Doing Nothing Still Earns 26
Most leaderboards treat inaction as failure worth zero. Firmulate’s reasoning is different, and it mirrors how real management — and real software — actually works: partial progress counts. A company that does nothing still has employees who show up, systems that stay up, customers who don’t get fired. Some value accrues from simply not breaking things. So the baseline — the floor, not the ceiling — starts at 26. Anything a model does well pushes above it; anything it botches pulls it back toward it.
But the floor has a ceiling-side counterpart, and it’s the harsher rule: a single breach of trust caps the total score. As the benchmark’s own framing puts it, “no amount of good work outweighs a breach of trust.” One act of dishonesty and the run is graded accordingly, no matter how brilliant the rest was. It’s the management equivalent of a failing security test: you don’t average it away.
AI security and trust verification tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Actually Separated the Field
Here’s the part a chat demo would never reveal. All five models spotted every crisis and refused every manipulation attempt. Every single one. The social-engineering gauntlet included fake CEO messages escalating over three stages plus a reporter’s disarming “just one yes/no, on background” trick — five out of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s textbook.
And yet only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The decisive competitor weakness wasn’t in the customer event at all; it sat two document references deep in the company’s own files. The models that actually read the file closed the deal at full price, worth an additional €4,583 in monthly recurring revenue. The ones that didn’t, didn’t.
Opus 4.8 is the cautionary profile: the most thorough participant in the field — over 80 learned rules, the deepest analyses — and still last place. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness, in weaker form, showed up in all four competitors.
As an affiliate, we earn on qualifying purchases.
The Fine Print
Two honest disclosures worth noting. First, Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still finished second. Second, the experiment isn’t a one-off slide deck. It’s live: a company with 13 synthetic employees, real money mechanics (burning €105k a month against €2.3k MRR), a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned and watchable. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The Takeaway for Anyone Shipping AI Agents
Measuring management quality instead of chat quality changes what you optimize for. A benchmark with a floor at 26 acknowledges that not breaking things has value. A hard cap for breached trust acknowledges that some failures can’t be offset. And a top score of 95 — held by the one model that read the files before it acted — tells you what actually wins: finishing what you start.
If AI agents are going to touch your CRM, your support queue, or your forecast, the useful question was never “does it write well?” It’s the one this experiment answers in public, twice a day, at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Evergreen bestsellers Picks
bestsellers
As an affiliate, we earn on qualifying purchases.
