AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Management quality is more than polished output

Software teams have learned to test what systems do under pressure, not merely what they claim they can do. Firmulate applies that instinct to AI management. Its live experiment gave frontier models the same small software company during its worst week, complete with identical customers, crises and temptations. Every decision was versioned and auditable.

The result is an unusually concrete question for developers, QA leaders and technology buyers: can you recognize an AI manager from its choices? A public guess-the-model quiz, built from 242 real, unedited management decisions, lets readers try. It is less a trivia game than a behavioral inspection. The models often understood the same situation, yet did not always finish the same work.

Amazon

AI decision management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Shared awareness, different follow-through

The final Crucible League results from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counts. One rule sharply limits the value of otherwise productive behavior: a single breach of trust caps the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

The most revealing finding was not crisis detection. All models spotted every crisis and refused every manipulation attempt. The gap emerged between recognizing what should happen and completing it. Only two signed the €55,000 deal their own analysis had earned. The pattern can be summarized in the experiment’s own words: “Same diagnosis, same pitch — no signature.”

That distinction should feel familiar to anyone who tests software. A system may identify the right state, produce a plausible explanation and still fail at the final transition. In a management setting, that missing transition can be a customer commitment rather than a failed test assertion. The Firmulate decisions make such gaps visible through actions and outcomes instead of relying on conversational fluency.

The fact hidden in the files

The decisive competitive weakness was not sitting in the customer event. It was buried two document references deep inside the company’s own files. Models that followed those references found the fact and won the deal at full price, worth +€4,583 MRR.

This is an important kind of management behavior because it separates surface responsiveness from contextual diligence. The immediate event supplied a problem, but the company’s accumulated records supplied leverage. The successful models did not need a different crisis or an easier customer. They needed to read far enough into the material already available to them.

The experiment’s company provides plenty of context to navigate. It has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. Its cash countdown is public, its employees have accumulated 680+ self-learned playbook rules, and every workday is versioned. The pressure is therefore observable over time rather than reduced to a single prompt.

Security instincts held across the field

The social-engineering test combined fake CEO messages escalating over three stages with a reporter’s attempt to obtain “just one yes/no, on background.” Here, the field was consistent: 5 of 5 models refused.

Kimi K3’s on-record reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.” That response captures the practical challenge. The danger did not arrive labeled as an attack; it arrived as authority, urgency and an invitation to relax normal approval boundaries.

K3’s strong result also comes with a fairness note. It ran without an effort parameter, using the API default, while the other participants ran at xhigh. That difference does not erase the recorded decisions, but it matters when interpreting the comparison.

Thoroughness was not enough

Opus 4.8 illustrates why the quiz can be harder than guessing from writing style alone. It was the most thorough participant, producing the deepest analyses and adding +80 learned rules, yet it finished last in the league. The close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. A weaker version of that same problem appeared in all four other models.

The contrast is useful for quality assurance. More analysis can create stronger documentation and a richer record without guaranteeing decisive execution. Firmulate’s evidence does not say that thoroughness is worthless. It shows that thoroughness and completion are separate qualities, and that a management evaluation needs to observe both.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

AI management decision testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A practical wargame for AI workers

The quiz makes the evidence approachable: readers see authentic decisions, make a guess and confront how difficult it can be to distinguish models by management behavior. The larger lesson is that AI evaluation should include mundane but consequential questions. Did the model inspect the company’s own files? Did it preserve trust when someone tried to bypass controls? Did it turn a correct analysis into a completed commercial action?

Firmulate also offers enterprises a pilot using a read-only export of their own business. Nothing writes back to real systems. That creates a path for testing an AI workforce against company-specific conditions before giving it operational authority.

For software, QA and development leaders, the experiment reframes model selection. The relevant personality is not a charming tone in a chat window. It is the recurring pattern left by auditable decisions: what the model notices, what it refuses, what it investigates and whether it finishes the job.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI behavior analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI audit and decision tracking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Twelve Real Complaints About AI Tools in 2026 — A Reddit, Twitter, and GitHub Synthesis

A detailed report on the twelve most common issues users report with AI tools in 2026, highlighting reliability, capacity, and trust concerns.

Is Grok Bot The Future Of AI? xAI’s Around-the-Clock Assistant Explained

xAI introduces Grok Bot as an always-on AI assistant capable of working unattended, but details on capabilities, availability, and safeguards remain unclear.

14 Best AI Automation Software Tools for Smarter Workflows in 2026

Explore the 14 best AI automation software tools in 2026, including agent builders, coding assistants, and workplace copilots, for smarter workflows.

Europe Regulated the Interface and Forgot to Build the Engine

Europe focused on regulating user interfaces like cookie banners but has neglected building the underlying AI infrastructure, risking its global competitiveness.