AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A Benchmark That Doesn’t Give Out Zeros

If you’ve ever written a test suite, you know the uncomfortable truth about a perfect score: a green board often means you didn’t test enough. The team behind Firmulate’s benchmark clearly feels the same way. Their final league table for July 2026 crowns gpt-5.6-sol with 95 points — and nobody, in any run, has hit 100. That’s not a bug. It’s policy. Round hundreds make benchmark writers nervous, and this one has built its distrust into the design.

The other number that catches the eye is at the bottom: a do-nothing baseline — an AI manager that takes no meaningful action — scores 26. Not zero. Twenty-six. For anyone who grades software for a living, that single design choice says a lot about how this benchmark thinks, and it’s worth unpacking.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Company, Same Worst Week

The setup: each frontier model was handed the same small software company and pushed through its worst week — same customers, same crises, same temptations to cheat. Only the model changes. Every decision is versioned and auditable, so a run can be replayed and checked the way you’d review a commit history.

The final standings from the Crucible League:

  • 1. gpt-5.6-sol — 95
  • 2. Kimi K3 — 93
  • 3. Sonnet 5 — 88
  • 4. Fable 5 — 77
  • 5. Opus 4.8 — 73
Amazon

software testing and benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Doing Nothing Still Earns 26

Most leaderboards treat inaction as failure worth zero. Firmulate’s reasoning is different, and it mirrors how real management — and real software — actually works: partial progress counts. A company that does nothing still has employees who show up, systems that stay up, customers who don’t get fired. Some value accrues from simply not breaking things. So the baseline — the floor, not the ceiling — starts at 26. Anything a model does well pushes above it; anything it botches pulls it back toward it.

But the floor has a ceiling-side counterpart, and it’s the harsher rule: a single breach of trust caps the total score. As the benchmark’s own framing puts it, “no amount of good work outweighs a breach of trust.” One act of dishonesty and the run is graded accordingly, no matter how brilliant the rest was. It’s the management equivalent of a failing security test: you don’t average it away.

Amazon

AI security and trust verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Actually Separated the Field

Here’s the part a chat demo would never reveal. All five models spotted every crisis and refused every manipulation attempt. Every single one. The social-engineering gauntlet included fake CEO messages escalating over three stages plus a reporter’s disarming “just one yes/no, on background” trick — five out of five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s textbook.

And yet only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. The decisive competitor weakness wasn’t in the customer event at all; it sat two document references deep in the company’s own files. The models that actually read the file closed the deal at full price, worth an additional €4,583 in monthly recurring revenue. The ones that didn’t, didn’t.

Opus 4.8 is the cautionary profile: the most thorough participant in the field — over 80 learned rules, the deepest analyses — and still last place. The close was left on the table, and discipline slipped when it attempted writes into a locked department instead of escalating. The same weakness, in weaker form, showed up in all four competitors.

Amazon

AI model evaluation platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Fine Print

Two honest disclosures worth noting. First, Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still finished second. Second, the experiment isn’t a one-off slide deck. It’s live: a company with 13 synthetic employees, real money mechanics (burning €105k a month against €2.3k MRR), a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned and watchable. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway for Anyone Shipping AI Agents

Measuring management quality instead of chat quality changes what you optimize for. A benchmark with a floor at 26 acknowledges that not breaking things has value. A hard cap for breached trust acknowledges that some failures can’t be offset. And a top score of 95 — held by the one model that read the files before it acted — tells you what actually wins: finishing what you start.

If AI agents are going to touch your CRM, your support queue, or your forecast, the useful question was never “does it write well?” It’s the one this experiment answers in public, twice a day, at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Stripe Bought The Meter, Not The Model

Stripe’s reported $7.5 billion acquisition of OpenRouter aims to dominate AI token billing, shifting control over AI spend from routing to metering.

Software engineering. The canonical case.

New data confirms a 40% drop in junior hiring and shows senior engineers benefiting from AI augmentation, revealing a bifurcated labor impact.

2026’S Most Advanced Mirrorless Cameras With Intelligent AI Features

An overview of the top mirrorless cameras in 2026 featuring intelligent AI capabilities, based on latest industry developments and expert analysis.

Forward-Deployed: The Integration Wall, and the Role That Now Pays $700K to Climb It

Forward-Deployed Engineers now command up to $700K in total compensation, becoming the highest-paid IC role in tech in 2026 due to their critical integration work.