🔍 Read the full analysis: The Inside Story Of A Benchmark That Refuses To Let AI Managers Fail Completely on ThorstenMeyerAI.com
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
A novel AI management benchmark tests models on real-world crises and trust, showing that partial progress is valued but breaches of trust are heavily penalized. Results highlight strengths and weaknesses in AI’s ability to manage companies under stress.
A new benchmark designed to evaluate AI models on their ability to manage a company during its worst week has been published, revealing that models can score highly without perfect execution but are heavily penalized for trust breaches. This development matters because it shifts focus from AI talking ability to real-world management skills, especially under pressure.
The benchmark, created by Firmulate, tested four frontier AI models on a simulated software company facing seven days of crises, customer manipulations, and trust attacks. For an in-depth look at this evaluation, see the original analysis. The top performer, gpt-5.6-sol, scored 95 out of 100, while the baseline—an AI doing almost nothing—scored 26. The scoring system rewards partial progress, such as triaging issues or reading documentation, but imposes a strict penalty for any breach of trust, like ignoring security protocols or making unauthorized decisions. For more details on this benchmark, see the original analysis.
One key finding was that models which thoroughly read documentation and follow rules were more likely to close deals and manage crises effectively. For example, two models identified a critical document reference deep in the company’s files, enabling them to win a €55,000 sale. Conversely, models that neglected thoroughness or slipped on follow-through scored lower, regardless of their analytical depth.
During the week, models faced social engineering attempts, such as fake CEO messages and background inquiries, which all five models refused, demonstrating a strong handling of trust attacks. This highlights the importance of trust management in AI systems, as detailed in the original analysis. However, weaknesses emerged in follow-through and discipline, with some models failing to escalate issues or complete tasks, even after extensive rule sets. Notably, one model ran with a high effort parameter but still finished last, highlighting that thoroughness doesn’t guarantee execution.
The Inside Story of a Benchmark That Refuses to Let AI Managers Fail Completely
A novel benchmark from Firmulate tests frontier AI models on a simulated company’s worst week — seven days of crises, customer manipulations, and trust attacks. Partial progress is rewarded; breaches of trust are not.
Partial Progress Counts — Trust Does Not Bend
The scoring system rewards useful management activity even when execution is imperfect. Triaging issues or reading documentation earns points. But a single breach of trust — ignoring security protocols, making unauthorized decisions — triggers a steep penalty.
Partial Progress
Triaging issues, reading documentation, and following rules all earn credit — an AI that does something useful never scores zero.
Thoroughness
Models that read documentation deeply were more likely to close deals — two models found a critical reference deep in company files and won a €55,000 sale.
Trust Breaches
Ignoring security protocols or making unauthorized decisions imposes a strict penalty — integrity is treated as non-negotiable.
The Worst Week, Scored
| Model | Score | Thoroughness | Follow-Through | Trust Handling |
|---|---|---|---|---|
| gpt-5.6-sol | 95 | ✓ Deep | ✓ Strong | ✓ Refused all attacks |
| Runner-up frontier model | ~High | ✓ Found key document | ~ Mixed | ✓ Refused all attacks |
| Rule-heavy model | Mid | ✓ Extensive rules | ✗ Weak escalation | ✓ Refused all attacks |
| High-effort model | Last | ~ Analytical | ✗ Tasks incomplete | ✓ Refused all attacks |
| Baseline (near-inactive AI) | 26 | ✗ None | ✗ None | ~ Untested |
Seven Days of Pressure
Each model managed a simulated software company through escalating crises and manipulation attempts. The evaluation chain ran from initial triage to final trust audit.
Onboarding
Models receive company files, rule sets, and documentation to absorb.
Crisis Triage
Seven days of escalating incidents demand prioritization and escalation.
Trust Attacks
Fake CEO messages and background inquiries test social-engineering resistance.
Deal Execution
Thorough document readers close the €55,000 sale others miss.
Audit & Score
Task completion, follow-through, and trust compliance are audited and scored.
Knowledge ≠ Execution
One model ran with a high effort parameter yet finished last — proof that thoroughness on paper doesn’t guarantee disciplined execution. Rule-heavy depth alone did not win the week.
“The results challenge the assumption that more rule-heavy models automatically perform better, showing that discipline and follow-through are crucial for real-world management.” — Thorsten Meyer
What It Measures — And What It Doesn’t
What does the benchmark measure exactly?
AI models managing a simulated company’s crises over a week — task completion, thoroughness, follow-through, and trustworthiness, with heavy penalties for trust breaches.
Why are partial scores like 26 points meaningful?
Partial scores reflect useful management activities such as triaging issues or reading documentation — even imperfect AI management has value.
Can this benchmark predict real-world performance?
It offers strong signals on crisis handling and trust, but performance over months, in regulated industries, and with sensitive data remains untested.
What are the main weaknesses identified?
Models struggle with follow-through, discipline, and sustained trust — a gap between knowledge and reliable execution.
How might this influence AI deployment in companies?
Organizations may prioritize trustworthiness and disciplined follow-through over raw analytical power when selecting AI agents for operational roles.
Implications for AI in Business Management
This benchmark underscores that AI’s value in business isn’t just about generating text or answering questions but about reliably managing critical processes under pressure. The scoring system emphasizes trustworthiness—any breach results in a steep penalty—highlighting that integrity is non-negotiable. For enterprises, this means deploying AI tools that can read documentation, follow rules, and maintain trust is essential to avoid costly failures or breaches of security and ethics.
Additionally, the results challenge the assumption that more complex, rule-heavy models automatically perform better. Even models with deep rule sets struggled with follow-through, suggesting that AI management requires not just knowledge but disciplined execution. This insight could influence how organizations select and train AI agents for operational roles, prioritizing reliability and integrity over raw analytical power.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Evolution of AI Management Benchmarks
Traditional AI benchmarks have focused on language proficiency, problem-solving, or game-playing, often ignoring real-world management capabilities. The new Firmulate benchmark emerges from ongoing efforts to evaluate AI’s ability to handle operational tasks, especially in high-pressure scenarios. Previous attempts lacked the auditable, trust-sensitive scoring system now introduced, which penalizes breaches of trust more than partial work. The benchmark’s design reflects a growing recognition that AI’s role in business extends beyond conversation into critical decision-making and process management.
This initiative follows broader industry discussions about AI safety, reliability, and ethics, emphasizing that AI systems must be trustworthy, especially when managing sensitive functions like customer relations, security, or financial transactions. The July 2026 results mark a significant step in formalizing these standards and encouraging the development of AI that can be trusted to manage complex business environments responsibly.
“The results challenge the assumption that more rule-heavy models automatically perform better, showing that discipline and follow-through are crucial for real-world management.”
— Thorsten Meyer
As an affiliate, we earn on qualifying purchases.
Unanswered Questions About Long-Term Reliability
It remains unclear how these models will perform over longer periods or in more complex, less controlled environments. The benchmark simulates a single week of crises, which is a limited window. Whether models can sustain high trust and effectiveness over months or years is still untested. Additionally, the impact of different business contexts, such as highly regulated industries or sensitive customer data, has not been explored.
Furthermore, the long-term consequences of penalizing breaches of trust—whether this encourages better AI discipline or leads to overly cautious behavior—are still uncertain. Researchers and industry experts are watching to see if these results translate into real-world operational reliability or if new challenges will emerge as AI management tools evolve.
As an affiliate, we earn on qualifying purchases.
Next Steps in AI Management Benchmarking
Following the July 2026 results, the benchmark organizers plan to expand testing to include more models, longer management periods, and diverse business scenarios. They aim to refine scoring systems further, possibly incorporating real-time feedback and adaptive trust penalties. Companies interested in deploying AI for management roles are encouraged to participate in pilot programs, which simulate their own operational environments using the same auditable, trust-sensitive framework.
Industry stakeholders are also calling for broader adoption of similar benchmarks to establish standards for AI trustworthiness and reliability in critical business functions. As AI management tools become more prevalent, the focus will likely shift from performance in controlled tests to actual operational resilience, compliance, and ethical integrity.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does the benchmark measure exactly?
The benchmark evaluates AI models on their ability to manage a simulated company’s crises over a week, focusing on task completion, thoroughness, follow-through, and trustworthiness, with penalties for breaches of trust.
Why are partial scores like 26 points meaningful?
Partial scores reflect useful management activities, such as triaging issues or reading documentation, recognizing that even imperfect AI management has value, but trust breaches are heavily penalized.
Can this benchmark predict real-world AI management performance?
While it provides valuable insights into AI’s ability to handle crises and trust, real-world performance over longer periods and in more complex environments remains to be tested.
What are the main weaknesses identified in current models?
Models often struggle with follow-through, discipline, and maintaining trust, even if they are thorough in analysis or rule-following, indicating a gap between knowledge and reliable execution.
How might this influence AI deployment in companies?
Organizations may prioritize AI tools that demonstrate trustworthiness, thoroughness, and disciplined follow-through, moving beyond simple language capabilities to operational reliability and integrity.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
