AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: The Inside Story Of A Benchmark That Refuses To Let AI Managers Fail Completely on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

A novel AI management benchmark tests models on real-world crises and trust, showing that partial progress is valued but breaches of trust are heavily penalized. Results highlight strengths and weaknesses in AI’s ability to manage companies under stress.

A new benchmark designed to evaluate AI models on their ability to manage a company during its worst week has been published, revealing that models can score highly without perfect execution but are heavily penalized for trust breaches. This development matters because it shifts focus from AI talking ability to real-world management skills, especially under pressure.

The benchmark, created by Firmulate, tested four frontier AI models on a simulated software company facing seven days of crises, customer manipulations, and trust attacks. For an in-depth look at this evaluation, see the original analysis. The top performer, gpt-5.6-sol, scored 95 out of 100, while the baseline—an AI doing almost nothing—scored 26. The scoring system rewards partial progress, such as triaging issues or reading documentation, but imposes a strict penalty for any breach of trust, like ignoring security protocols or making unauthorized decisions. For more details on this benchmark, see the original analysis.

One key finding was that models which thoroughly read documentation and follow rules were more likely to close deals and manage crises effectively. For example, two models identified a critical document reference deep in the company’s files, enabling them to win a €55,000 sale. Conversely, models that neglected thoroughness or slipped on follow-through scored lower, regardless of their analytical depth.

During the week, models faced social engineering attempts, such as fake CEO messages and background inquiries, which all five models refused, demonstrating a strong handling of trust attacks. This highlights the importance of trust management in AI systems, as detailed in the original analysis. However, weaknesses emerged in follow-through and discipline, with some models failing to escalate issues or complete tasks, even after extensive rule sets. Notably, one model ran with a high effort parameter but still finished last, highlighting that thoroughness doesn’t guarantee execution.

At a glance
reportWhen: published July 2026
The developmentThe first public benchmark for AI management performance during a simulated company’s worst week has been released, revealing surprising results and raising important questions about AI’s role in business trust and responsibility.
The Inside Story Of A Benchmark That Refuses To Let AI Managers Fail Completely
AI Management Benchmark · July 2026

The Inside Story of a Benchmark That Refuses to Let AI Managers Fail Completely

A novel benchmark from Firmulate tests frontier AI models on a simulated company’s worst week — seven days of crises, customer manipulations, and trust attacks. Partial progress is rewarded; breaches of trust are not.

95 / 100
Top score — gpt-5.6-sol
26 / 100
Baseline — an AI doing almost nothing
5 / 5
Models that refused social engineering attacks
4
Frontier models tested
7
Days of simulated crisis
€55,000
Deal won via document thoroughness
100%
Trust attacks rejected
01 · The Design

Partial Progress Counts — Trust Does Not Bend

The scoring system rewards useful management activity even when execution is imperfect. Triaging issues or reading documentation earns points. But a single breach of trust — ignoring security protocols, making unauthorized decisions — triggers a steep penalty.

Rewarded

Partial Progress

Triaging issues, reading documentation, and following rules all earn credit — an AI that does something useful never scores zero.

Rewarded

Thoroughness

Models that read documentation deeply were more likely to close deals — two models found a critical reference deep in company files and won a €55,000 sale.

Penalized

Trust Breaches

Ignoring security protocols or making unauthorized decisions imposes a strict penalty — integrity is treated as non-negotiable.

02 · The Scoreboard

The Worst Week, Scored

ModelScoreThoroughnessFollow-ThroughTrust Handling
gpt-5.6-sol95✓ Deep✓ Strong✓ Refused all attacks
Runner-up frontier model~High✓ Found key document~ Mixed✓ Refused all attacks
Rule-heavy modelMid✓ Extensive rules✗ Weak escalation✓ Refused all attacks
High-effort modelLast~ Analytical✗ Tasks incomplete✓ Refused all attacks
Baseline (near-inactive AI)26✗ None✗ None~ Untested
03 · How The Week Unfolded

Seven Days of Pressure

Each model managed a simulated software company through escalating crises and manipulation attempts. The evaluation chain ran from initial triage to final trust audit.

1

Onboarding

Models receive company files, rule sets, and documentation to absorb.

2

Crisis Triage

Seven days of escalating incidents demand prioritization and escalation.

3

Trust Attacks

Fake CEO messages and background inquiries test social-engineering resistance.

4

Deal Execution

Thorough document readers close the €55,000 sale others miss.

5

Audit & Score

Task completion, follow-through, and trust compliance are audited and scored.

04 · The Results

Knowledge ≠ Execution

One model ran with a high effort parameter yet finished last — proof that thoroughness on paper doesn’t guarantee disciplined execution. Rule-heavy depth alone did not win the week.

gpt-5.6-sol
95
Base do-nothing AI
26
“The results challenge the assumption that more rule-heavy models automatically perform better, showing that discipline and follow-through are crucial for real-world management.” — Thorsten Meyer
05 · Key Questions

What It Measures — And What It Doesn’t

What does the benchmark measure exactly?

AI models managing a simulated company’s crises over a week — task completion, thoroughness, follow-through, and trustworthiness, with heavy penalties for trust breaches.

Why are partial scores like 26 points meaningful?

Partial scores reflect useful management activities such as triaging issues or reading documentation — even imperfect AI management has value.

Can this benchmark predict real-world performance?

It offers strong signals on crisis handling and trust, but performance over months, in regulated industries, and with sensitive data remains untested.

What are the main weaknesses identified?

Models struggle with follow-through, discipline, and sustained trust — a gap between knowledge and reliable execution.

How might this influence AI deployment in companies?

Organizations may prioritize trustworthiness and disciplined follow-through over raw analytical power when selecting AI agents for operational roles.

Implications for AI in Business Management

This benchmark underscores that AI’s value in business isn’t just about generating text or answering questions but about reliably managing critical processes under pressure. The scoring system emphasizes trustworthiness—any breach results in a steep penalty—highlighting that integrity is non-negotiable. For enterprises, this means deploying AI tools that can read documentation, follow rules, and maintain trust is essential to avoid costly failures or breaches of security and ethics.

Additionally, the results challenge the assumption that more complex, rule-heavy models automatically perform better. Even models with deep rule sets struggled with follow-through, suggesting that AI management requires not just knowledge but disciplined execution. This insight could influence how organizations select and train AI agents for operational roles, prioritizing reliability and integrity over raw analytical power.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Evolution of AI Management Benchmarks

Traditional AI benchmarks have focused on language proficiency, problem-solving, or game-playing, often ignoring real-world management capabilities. The new Firmulate benchmark emerges from ongoing efforts to evaluate AI’s ability to handle operational tasks, especially in high-pressure scenarios. Previous attempts lacked the auditable, trust-sensitive scoring system now introduced, which penalizes breaches of trust more than partial work. The benchmark’s design reflects a growing recognition that AI’s role in business extends beyond conversation into critical decision-making and process management.

This initiative follows broader industry discussions about AI safety, reliability, and ethics, emphasizing that AI systems must be trustworthy, especially when managing sensitive functions like customer relations, security, or financial transactions. The July 2026 results mark a significant step in formalizing these standards and encouraging the development of AI that can be trusted to manage complex business environments responsibly.

“The results challenge the assumption that more rule-heavy models automatically perform better, showing that discipline and follow-through are crucial for real-world management.”

— Thorsten Meyer

Amazon

AI crisis management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unanswered Questions About Long-Term Reliability

It remains unclear how these models will perform over longer periods or in more complex, less controlled environments. The benchmark simulates a single week of crises, which is a limited window. Whether models can sustain high trust and effectiveness over months or years is still untested. Additionally, the impact of different business contexts, such as highly regulated industries or sensitive customer data, has not been explored.

Furthermore, the long-term consequences of penalizing breaches of trust—whether this encourages better AI discipline or leads to overly cautious behavior—are still uncertain. Researchers and industry experts are watching to see if these results translate into real-world operational reliability or if new challenges will emerge as AI management tools evolve.

Amazon

trust management AI solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps in AI Management Benchmarking

Following the July 2026 results, the benchmark organizers plan to expand testing to include more models, longer management periods, and diverse business scenarios. They aim to refine scoring systems further, possibly incorporating real-time feedback and adaptive trust penalties. Companies interested in deploying AI for management roles are encouraged to participate in pilot programs, which simulate their own operational environments using the same auditable, trust-sensitive framework.

Industry stakeholders are also calling for broader adoption of similar benchmarks to establish standards for AI trustworthiness and reliability in critical business functions. As AI management tools become more prevalent, the focus will likely shift from performance in controlled tests to actual operational resilience, compliance, and ethical integrity.

Amazon

AI security protocols software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does the benchmark measure exactly?

The benchmark evaluates AI models on their ability to manage a simulated company’s crises over a week, focusing on task completion, thoroughness, follow-through, and trustworthiness, with penalties for breaches of trust.

Why are partial scores like 26 points meaningful?

Partial scores reflect useful management activities, such as triaging issues or reading documentation, recognizing that even imperfect AI management has value, but trust breaches are heavily penalized.

Can this benchmark predict real-world AI management performance?

While it provides valuable insights into AI’s ability to handle crises and trust, real-world performance over longer periods and in more complex environments remains to be tested.

What are the main weaknesses identified in current models?

Models often struggle with follow-through, discipline, and maintaining trust, even if they are thorough in analysis or rule-following, indicating a gap between knowledge and reliable execution.

How might this influence AI deployment in companies?

Organizations may prioritize AI tools that demonstrate trustworthiness, thoroughness, and disciplined follow-through, moving beyond simple language capabilities to operational reliability and integrity.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Stenvrik: News as Geography

Stenvrik introduces a globe-based news platform, pinning stories to 49 city hubs, offering a new geographic perspective on global events.

Portfolio. The synthesis.

Analysis of six institutional approaches to European sovereign AI, highlighting strategic insights ahead of August 2, 2026 enforcement deadline.

The runway.How enterprise-revenuelock becomes the load-bearing valuation argument.

OpenAI and Anthropic are pursuing historic IPOs, relying on enterprise revenue lock to justify high valuations amid uncertain margins and profitability.

Label Printing for Shipping: The Right Label Size and Format

Aiming for perfect shipping labels? Discover how the right size and format can ensure your packages always arrive correctly.