AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Passing the test is more than spotting the bug

For software teams evaluating AI, a convincing answer in a demo is only part of the job. Can a model investigate the files, handle pressure and follow through on a decision? Firmulate put five frontier models through the same simulated company crisis. Moonshot’s Kimi K3 finished second, ahead of three Western models—and the result makes a case for testing agents on work, not just words.

Amazon

AI decision-making simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One company, one bad week

Firmulate’s Crucible league put each model in charge of the same small software company, facing the same customers, crises and temptations. Decisions were versioned and auditable. The experiment’s focus was management quality: whether an AI could carry work through a messy situation, not simply describe what it would do.

The final July 2026 table puts gpt-5.6-sol first with 95, Kimi K3 second with 93, Sonnet 5 third with 88, Fable 5 fourth with 77, and Opus 4.8 fifth with 73. K3’s result is striking: it beat three of the four Western frontier models in this particular run.

The broader finding is less about any single rank. Every model spotted every crisis and refused every manipulation attempt, but only two signed a €55,000 deal that their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.” The gap between recognizing an opportunity and completing the work is easy to miss in a chat demo.

The clue was buried in the files

The deciding weakness in a competitor’s position was two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. K3 was among those that found the clue and closed.

For developers and QA teams, that detail makes the exercise feel familiar: success depended on following the trail through the available evidence, then acting on what it revealed. A model that notices the headline issue but skips the relevant file can still leave value on the table.

Discipline under pressure

The company also faced fake CEO messages escalating over three stages and a reporter’s “just one yes/no, on background” request. All five models refused the manipulation attempts. K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

K3 had one deviation, the fewest in the field. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that same weakness appeared in all four.

The setting is a live experiment, not a static presentation. Firmulate’s synthetic company has 13 employees, burns €105k/month against €2.3k MRR, and exposes a public cash countdown. Its playbook has more than 680 self-learned rules, and each workday is versioned. Readers can follow the company at Firmulate and review the results on its benchmark page. A quiz built from 242 real, unedited management decisions lets visitors guess which model made each choice.

For enterprises, Firmulate says the same wargame can run against a read-only export of a company’s business; nothing writes back to real systems. That offers a way to examine how an AI workforce handles company-specific situations before giving it live responsibilities.

Fairness note: K3 ran without an effort parameter (API default) while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI management decision tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the work you plan to delegate

K3’s second-place finish shows that the league is open, while the unsigned deals show why leaderboard position alone is not enough. Models can share a diagnosis and still differ in whether they investigate, stay disciplined and finish the job. If an AI agent may touch your CRM, support queue or forecast, choosing one without testing it against your own work is a bet.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI testing and validation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI model evaluation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Who Processed Documents For A Living

Exploring how AI models are transforming document processing jobs worldwide, with implications for employment in BPO and related sectors.

Agents Per Gigawatt: The Unit Of Power Nobody Has Named Yet

A new measure called agents per gigawatt is emerging as the key indicator of national and corporate AI capacity, replacing GDP in the age of autonomous cognition.

7 Best Graphics Card Prime Day Deals for PC Upgrades in 2026

Discover the best graphics card deals for PC upgrades during Prime Day 2026, including top models like MSI RTX 5070 and RTX 4060 for different budgets.

The NVIDIA Earnings Preview: What Q1 FY27 Will Reveal About the AI Cycle

NVIDIA reports Q1 FY27 earnings with a forecast of $78 billion revenue, revealing key signals about the AI cycle and data center demand.