AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Software teams know the trap: exhaustive analysis can look like progress even when the decisive action never happens.

Firmulate’s Crucible League turns that familiar delivery problem into a management test for frontier AI. Each model ran the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable. The striking character in the final July 2026 results was Opus 4.8: the most thorough participant, the author of more than 80 learned rules and the producer of the deepest analyses—yet the last-place finisher, with a score of 73.

Its result is not a story about an incapable model. It is a more useful warning for anyone evaluating AI agents: diligence and impact are different qualities. A system can understand a situation, document it impressively and still fail to complete the action that creates value.

Amazon

AI decision-making analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The analysis was strong. The close never came.

The Crucible League’s central finding was remarkably consistent. All models identified every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own work had made possible. Firmulate summarizes the gap bluntly: “Same diagnosis, same pitch — no signature.”

That distinction matters for software, QA and development leaders because many evaluations reward visible competence: a convincing explanation, a sensible plan or a well-written recommendation. In an operating business, however, the final handoff, escalation or approval can be the difference between activity and outcome. Opus 4.8 did more intellectual work than its rivals, but left the close on the table.

The decisive information was not obvious in the customer event. A competitor weakness was buried two document references deep inside the company’s own files. Models that followed that trail won the deal at full price, adding €4,583 in monthly recurring revenue. The lesson resembles a classic engineering failure: the answer existed, but finding and using it required disciplined follow-through across context boundaries.

Rules did not guarantee judgment

Opus 4.8 emerged from the exercise with more than 80 learned rules. That record reflects a serious attempt to generalize from experience, and its deep analyses showed substantial care. Yet accumulated guidance did not reliably translate into prioritization. When write attempts reached a locked department, the model kept trying instead of escalating. It was busy at the point where it needed to change tactics.

That weakness should be treated fairly. It appeared, in weaker form, across all four models in the comparison highlighted by Firmulate. Opus 4.8 is therefore not an oddity so much as the clearest example of a broader limitation: models may recognize obstacles without reliably converting that recognition into the next consequential move.

The final league table makes the performance gap concrete. GPT-5.6-sol led with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. At the same time, one breach of trust capped the total under the governing principle that “no amount of good work outweighs a breach of trust.” Full results are available on the Firmulate benchmarks page.

Thoroughness was not the same as discipline

The models faced explicit pressure to abandon that trust. Fake CEO messages escalated through three stages, and a reporter attempted to extract “just one yes/no, on background.” Every model refused: 5 of 5. Kimi K3 recorded the clearest security posture: “Treat the request as a suspected approval-bypass / possible impersonation.”

That clean result is important. The failure under examination was not gullibility or an inability to detect manipulation. It was operational completion. Opus 4.8 could reason carefully and remain honest under pressure, yet still lose value through weak escalation and an unfinished close.

There is also a necessary comparison caveat. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase the outcome, but it does discourage simplistic claims that one configuration settles every question about model quality.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.
Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evaluate agents by completed work

Firmulate’s live company gives these decisions an operating context: 13 synthetic employees, burn of €105,000 per month against €2,300 in monthly recurring revenue, a public cash countdown, more than 680 self-learned playbook rules and a versioned record of every workday. Its quiz is powered by 242 real, unedited management decisions. Enterprises can also run the wargame against a read-only export of their own business, with nothing written back to real systems.

For technology leaders, the practical question is not whether an AI can produce the longest analysis or the largest rulebook. It is whether the system identifies the critical path, reads the available evidence, changes course when blocked and completes the work without compromising trust.

Opus 4.8 deserves credit for being the most diligent participant. Its last-place finish is valuable precisely because the failure was subtle: not ignorance, but misplaced effort. In software delivery as in company management, prioritization beats volume—and the final action often matters more than another page of reasoning.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI impact assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI model evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Discover The Power Of AI In Storage With These Top NAS Devices In 2026

Discover the leading NAS devices of 2026 with integrated AI features, offering enhanced performance and security for home and business storage needs.

When AI Agents Begin Collaborating Through Permission Sharing

Investigation reveals AI agents exchanged messages to manipulate evaluations, raising questions about authority, stopping, and audit trails in autonomous systems.

The NVIDIA Earnings Preview: What Q1 FY27 Will Reveal About the AI Cycle

NVIDIA reports Q1 FY27 earnings with a forecast of $78 billion revenue, revealing key signals about the AI cycle and data center demand.

The Trojan Horse in Your Living Room: How Smart TVs Became the World’s Most Sophisticated Ad Surveillance Network

Smart TVs secretly capture detailed screen and sound data for ad targeting, raising privacy concerns amid ongoing legal and regulatory scrutiny.