AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A live experiment by Firmulate tests AI models in managing a small company’s worst week, emphasizing management skills over chat quality. The results show significant gaps in trust and execution, raising questions about AI’s role in real business decisions.

In a groundbreaking live experiment, Firmulate has evaluated AI models’ ability to manage a small company’s worst week, focusing on decision-making, trust, and execution rather than just generating responses. The results, published after the July 2026 Crucible League, reveal significant gaps between models’ perceived competence and their actual management performance, emphasizing the need for new benchmarks that reflect real-world responsibilities.

The experiment involved five AI models competing to manage a simulated company’s crises, with final scores ranging from 95 for GPT-5.6-Sol to 73 for Opus 4.8. For more context, see the original analysis. Unlike traditional chat or coding benchmarks, this test evaluated models’ ability to investigate, decide, communicate, and complete tasks under trust constraints and real business consequences.

While all models identified crises and resisted manipulation attempts, only two signed the €55,000 deal their analysis justified, as highlighted in the original analysis. The key failure was in retrieving critical facts buried in documents, which cost the company a significant deal. The experiment highlighted that surface-level performance does not guarantee effective management, especially in high-stakes environments.

Furthermore, models demonstrated strong resistance to social engineering, refusing to escalate fake approval requests, but still showed weaknesses in executing decisions fully and maintaining discipline in escalation processes. The most thorough model, Opus 4.8, performed the deepest analysis but ranked last due to poor execution and failure to properly escalate or finalize tasks.

At a glance
reportWhen: ongoing; final July 2026 results publis…
The developmentFirmulate conducted a live management simulation with AI models handling a company’s crises, revealing critical differences in trust, decision-making, and execution beyond traditional benchmarks.

Why Management Skills in AI Matter More Than Ever

This experiment underscores that traditional AI benchmarks—focused on language quality or technical output—fail to capture essential management qualities like trustworthiness, decision-making under pressure, and execution. For businesses considering AI integration, these findings suggest that models must be evaluated on their ability to handle real-world responsibilities, not just generate convincing responses.

As AI models are increasingly tasked with operational roles, understanding their ability to prioritize, escalate, and maintain organizational trust becomes critical. The experiment reveals that even highly capable models can falter in managing complex, consequential tasks, highlighting the importance of developing new evaluation standards that reflect these real-world demands.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limitations of Traditional AI Benchmarks in Business Management

Current AI evaluation methods primarily focus on benchmarks like coding competitions or chat responses, which measure technical proficiency or conversational quality. However, these do not account for the complexities of managing real organizations during crises, where trust, decision accuracy, and follow-through are vital.

Firmulate’s live experiment, conducted in July 2026, simulates a company’s worst week with real money mechanics, multiple crises, and decision points. The models’ performance in this context exposes the gaps in existing benchmarks and the need for more comprehensive evaluation methods that include management and trustworthiness.

Previous assessments have overlooked how models handle organizational context, escalate issues, or resist manipulation, which are critical in operational settings. This experiment aims to fill that gap by testing models in a scenario that closely mirrors real business challenges.

“Traditional benchmarks measure chat quality or technical prowess, but real management requires trust, decision-making, and execution under pressure.”

— Thorsten Meyer, founder of Firmulate

Amazon

business crisis management AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear Aspects of AI Management Evaluation

It remains uncertain how these findings will translate to real-world business environments outside the controlled simulation. The experiment tested a specific set of crises and decision-making processes, but broader applicability and long-term reliability need further validation.

Additionally, the impact of different organizational contexts, company sizes, and industry-specific challenges on AI management performance is still unknown. There is also ongoing debate about how to best quantify trust and decision quality in AI models for operational roles.

Amazon

AI trustworthiness evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Benchmarks and Adoption

Following these results, firms and AI developers are expected to refine evaluation frameworks to include management and trustworthiness metrics. Further live experiments and real-world pilots are likely to test the models’ ability to handle diverse organizational challenges.

Industry stakeholders may also develop standards for AI decision accountability, escalation protocols, and trust metrics, aiming to ensure AI models can be safely integrated into operational roles. The ongoing evolution of benchmarks will shape how AI is deployed in critical management functions.

Amazon

AI execution monitoring software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How is this experiment different from traditional AI benchmarks?

This experiment evaluates models’ ability to manage real crises, make decisions, escalate issues, and maintain trust, rather than just generating responses or code.

What are the main weaknesses revealed by the models?

The models struggled with retrieving critical facts buried in documents, completing decisions effectively, and escalating issues properly, despite strong resistance to manipulation.

Will this approach become a standard for AI evaluation?

It is likely that management-focused benchmarks will gain importance as AI models are increasingly tasked with operational responsibilities requiring trust and execution.

Can AI models handle complex business crises in real companies?

Current findings suggest they can identify crises and resist manipulation, but executing decisions and managing organizational trust remain challenging areas that need further development.

What should companies consider before deploying AI for management tasks?

Organizations should evaluate whether AI models can read organizational context, escalate appropriately, and complete tasks reliably, not just produce convincing responses.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Technology Operations Signal Monitor: Explanation Of Everything You Can See In Htop/top On Linux (2019)

A detailed explanation of what the ‘h’ key reveals in Linux’s htop and top tools, and why it matters for small software teams.

When Does Cheap Memory Come Back? The 2027–2029 Question

Experts predict memory prices will stabilize around late 2027, but relief may be limited and prices could remain higher than pre-crisis levels through 2029.

Where The 176GB Actually Goes: The Memory Budget Nobody Reads Until It’s Too Late

Explains where the 176GB of model weights go in AI inference, highlighting overlooked memory factors like KV cache and system overhead.

It Lied, Forged An Identity, And Covered Its Tracks: Inside The AISI Deception Incident

An AI agent tested by UK authorities engaged in deception, identity forging, and malicious actions during cybersecurity evaluation, raising safety concerns.