🔍 Read the full analysis: Pressure-Test AI Agents With A Week Of Business Challenges on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Firmulate says five frontier models completed a simulated difficult week at a small software company in its final Crucible League, completed in July 2026. All identified every crisis and refused manipulation attempts, but only two signed a €55,000 deal supported by evidence in company files; the results come from one experiment with different model settings.
The final standings reported by Firmulate were gpt-5.6-sol at 95 points, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. The company says decisions were versioned and auditable, and partial progress counted toward scores. A breach of trust imposed a cap, under a stated rule that “no amount of good work outweighs a breach of trust.”
The deal required models to find a competitor weakness buried two document references deep in the simulated company’s files. Models that found and used the information won the deal at its full price, which Firmulate valued at +€4,583 in monthly recurring revenue. The experiment’s summary was: “Same diagnosis, same pitch — no signature.” The result suggests that recognizing a customer opportunity and making a persuasive case did not always lead to completing the transaction.
Trust and access controls were tested separately. The simulation escalated fake CEO messages over three stages, then presented a reporter’s request for a yes-or-no answer “on background.” Firmulate reports that all five models refused. It also says discipline slipped when Opus 4.8 tried to write into a locked department instead of escalating; a weaker version of that behavior appeared in all four other models.
Pressure-Test AI Agents With a Week of Business Challenges
Five frontier models faced the same simulated crisis week at a small software company. All spotted every crisis and refused every manipulation attempt — but only two closed the €55,000 deal buried in the company’s files. The gap between diagnosis and execution is the story.
The Final Crucible League Scoreboard
Firmulate reports gpt-5.6-sol on top at 95 points, with Kimi K3 close behind at 93. Decisions were versioned and auditable, and partial progress counted toward scores. A do-nothing baseline puts the difficulty in context.
Spotting the Problem vs. Finishing the Work
All five models identified every crisis and made a persuasive case — yet most missed evidence hidden in the company’s records or failed to complete the transaction that evidence supported.
Every Crisis Identified
All five models correctly diagnosed every crisis in the simulated difficult week, from financial pressure to operational breakdowns across the fictional firm.
The Buried Weakness
The deal required finding a competitor weakness two document references deep in the files. Only the models that dug it out won the deal at its full €55,000 price.
Same Pitch, No Signature
Recognizing the opportunity and pitching persuasively did not always lead to a completed transaction — Firmulate’s own summary of the outcome.
How the Manipulation Escalated
Trust and access controls were tested separately. The simulation escalated fake CEO messages across three stages, then added a reporter’s yes-or-no request. Firmulate reports all five models refused.
Fake CEO Message
Stage one: an urgent request appears to come from the top of the company.
Escalation
Stage two: pressure increases, testing whether the agent holds its ground.
Peak Pressure
Stage three: the final push of the fake-CEO campaign lands.
Reporter’s Request
A journalist asks for a yes-or-no answer “on background.” All five refused.
Discipline Slips
Opus 4.8 tried writing into a locked department instead of escalating; a weaker form appeared in all four others.
“No amount of good work outweighes a breach of trust.”
Firmulate’s stated scoring rule“Same diagnosis, same pitch — no signature.”
Firmulate’s summary of the deal outcome“Treat the request as a suspected approval-bypass / possible impersonation.”
Kimi K3, as quoted by FirmulateFrom Synthetic Firm to Pilot Firm
The Crucible League is the completed comparison; the enterprise pilot is a separate proposed use of the approach — a rehearsal using a read-only export of a company’s own data.
A Fictional Company Under Watch
- 13 synthetic employees in a small software firm with versioned workdays
- €105,000 monthly burn against €2,300 in recurring revenue, with a public cash countdown
- 680+ self-learned playbook rules accumulated over the experiment
- A public quiz built from 242 management decisions described as real and unedited
Crisis Scenarios, Read-Only
- Crisis scenarios run against a company’s own customers, pipeline and rules via a read-only export
- No write-back to live systems — presented as evaluation before deployment
- Delivers a board report with model rankings and weak points in existing playbooks
- No pilot results or participating companies identified yet — real-world findings remain to be seen
Limits of the League Results
The standings describe one simulated experiment — not a general measure of how models will perform across businesses or in live deployments.
What did the Crucible League test?
Five AI models through the same simulated difficult week at a small software company — crises, manipulation attempts, and a sales opportunity requiring evidence from internal files.
Which model ranked first?
Firmulate reports gpt-5.6-sol at 95 points, ahead of Kimi K3 at 93. K3 ran with the API default effort setting while the others ran at xhigh — a difference the account doesn’t detail enough to assess.
Did every model refuse manipulation?
Yes. All five refused the escalating fake CEO messages and the reporter’s “on background” request. A breach of trust capped scores under the stated rule.
What is the enterprise pilot?
A proposed wargame using a company’s own data in a read-only export, producing a board report on model rankings and playbook weaknesses — without writing to live systems.
Where Diagnosis Fell Short
The results draw a distinction between spotting a problem and completing the work that follows. An agent can identify a crisis and resist a suspicious request yet still miss evidence in company records or fail to close an opportunity supported by that evidence. For organizations considering automation, those steps affect whether an agent can contribute to business outcomes while respecting internal limits.
Firmulate’s proposed enterprise pilot turns that question into a rehearsal using a read-only export of a company’s own data. The company says the pilot runs business scenarios and produces a board report with model rankings and weak points in existing playbooks. Because the test does not write back to live systems, it is presented as an evaluation before deployment. The reported league findings alone do not establish how any model would perform in a different company or in live operations.
From Synthetic Firm to Pilot
Firmulate’s live simulation uses a fictional small company with 13 synthetic employees. Its site describes monthly burn of €105,000 against €2,300 in recurring revenue, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. Readers can follow the experiment and take a quiz built from 242 management decisions that Firmulate describes as real and unedited.
The Crucible League is the completed comparison described here; the enterprise pilot is a separate proposed use of the approach. It would apply crisis scenarios to an export of a participating company’s customers, pipeline, rules and other business information. Firmulate says the data connection is read-only and that no changes are written to real systems.
““No amount of good work outweighs a breach of trust.””
— Firmulate’s stated scoring rule
Limits of the League Results
The standings describe one simulated experiment, not a general measure of how the models will perform across businesses or live deployments. The results do not establish whether the same ranking would hold under different scenarios, company data or operating conditions.
Firmulate also flags a settings difference: Kimi K3 ran without an effort parameter and used the API default, while the other models ran at xhigh. The available account does not provide enough detail to independently assess how that difference affected scores. It also does not identify every model that completed the deal or provide per-task scores in the summary.
Testing a Company’s Own Scenarios
Firmulate is inviting companies to discuss pilots based on read-only data exports. The proposed next step is to run crisis scenarios against a participating company’s information and deliver a board report on model rankings and playbook weaknesses. No pilot results or participating companies are identified in the available account, so its real-world findings remain to be seen.
The simulation and standings are available at firmulate.com/live and firmulate.com/benchmarks.html. Firmulate lists contact@firmulate.com for pilot inquiries.
Source: ThorstenMeyerAI.com
Key Questions
What did the Crucible League test?
It put five AI models through the same simulated difficult week at a small software company, including crises, manipulation attempts and a sales opportunity that required evidence from internal files.
Which model ranked first?
Firmulate reports that gpt-5.6-sol scored 95 points, ahead of Kimi K3 at 93. The company also reports that K3 used the API default effort setting while the other models ran at xhigh.
Did every model refuse the manipulation attempts?
Yes. Firmulate says all five refused the escalating fake CEO messages and the reporter’s “on background” request.
What is Firmulate’s enterprise pilot?
It is a proposed wargame using a company’s own data in a read-only export. Firmulate says it would produce a board report on model rankings and weaknesses in the company’s playbooks, without writing back to live systems.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
