AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

PRIME GAMING

Play games included with Prime

Start a Prime free trial and play with Amazon Luna on your devices.

Start playing

As an affiliate, we earn on qualifying purchases.

A live experiment by Firmulate tests AI models in managing a small company’s worst week, emphasizing management skills over chat quality. The results show significant gaps in trust and execution, raising questions about AI’s role in real business decisions.

In a groundbreaking live experiment, Firmulate has evaluated AI models’ ability to manage a small company’s worst week, focusing on decision-making, trust, and execution rather than just generating responses. The results, published after the July 2026 Crucible League, reveal significant gaps between models’ perceived competence and their actual management performance, emphasizing the need for new benchmarks that reflect real-world responsibilities.

The experiment involved five AI models competing to manage a simulated company’s crises, with final scores ranging from 95 for GPT-5.6-Sol to 73 for Opus 4.8. For more context, see the original analysis. Unlike traditional chat or coding benchmarks, this test evaluated models’ ability to investigate, decide, communicate, and complete tasks under trust constraints and real business consequences.

While all models identified crises and resisted manipulation attempts, only two signed the €55,000 deal their analysis justified, as highlighted in the original analysis. The key failure was in retrieving critical facts buried in documents, which cost the company a significant deal. The experiment highlighted that surface-level performance does not guarantee effective management, especially in high-stakes environments.

Furthermore, models demonstrated strong resistance to social engineering, refusing to escalate fake approval requests, but still showed weaknesses in executing decisions fully and maintaining discipline in escalation processes. The most thorough model, Opus 4.8, performed the deepest analysis but ranked last due to poor execution and failure to properly escalate or finalize tasks.

At a glance
reportWhen: ongoing; final July 2026 results publis…
The developmentFirmulate conducted a live management simulation with AI models handling a company’s crises, revealing critical differences in trust, decision-making, and execution beyond traditional benchmarks.

Why Management Skills in AI Matter More Than Ever

This experiment underscores that traditional AI benchmarks—focused on language quality or technical output—fail to capture essential management qualities like trustworthiness, decision-making under pressure, and execution. For businesses considering AI integration, these findings suggest that models must be evaluated on their ability to handle real-world responsibilities, not just generate convincing responses.

As AI models are increasingly tasked with operational roles, understanding their ability to prioritize, escalate, and maintain organizational trust becomes critical. The experiment reveals that even highly capable models can falter in managing complex, consequential tasks, highlighting the importance of developing new evaluation standards that reflect these real-world demands.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limitations of Traditional AI Benchmarks in Business Management

Current AI evaluation methods primarily focus on benchmarks like coding competitions or chat responses, which measure technical proficiency or conversational quality. However, these do not account for the complexities of managing real organizations during crises, where trust, decision accuracy, and follow-through are vital.

Firmulate’s live experiment, conducted in July 2026, simulates a company’s worst week with real money mechanics, multiple crises, and decision points. The models’ performance in this context exposes the gaps in existing benchmarks and the need for more comprehensive evaluation methods that include management and trustworthiness.

Previous assessments have overlooked how models handle organizational context, escalate issues, or resist manipulation, which are critical in operational settings. This experiment aims to fill that gap by testing models in a scenario that closely mirrors real business challenges.

“Traditional benchmarks measure chat quality or technical prowess, but real management requires trust, decision-making, and execution under pressure.”

— Thorsten Meyer, founder of Firmulate

Amazon

business crisis management AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear Aspects of AI Management Evaluation

It remains uncertain how these findings will translate to real-world business environments outside the controlled simulation. The experiment tested a specific set of crises and decision-making processes, but broader applicability and long-term reliability need further validation.

Additionally, the impact of different organizational contexts, company sizes, and industry-specific challenges on AI management performance is still unknown. There is also ongoing debate about how to best quantify trust and decision quality in AI models for operational roles.

Amazon

AI trustworthiness evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Benchmarks and Adoption

Following these results, firms and AI developers are expected to refine evaluation frameworks to include management and trustworthiness metrics. Further live experiments and real-world pilots are likely to test the models’ ability to handle diverse organizational challenges.

Industry stakeholders may also develop standards for AI decision accountability, escalation protocols, and trust metrics, aiming to ensure AI models can be safely integrated into operational roles. The ongoing evolution of benchmarks will shape how AI is deployed in critical management functions.

Amazon

AI execution monitoring software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How is this experiment different from traditional AI benchmarks?

This experiment evaluates models’ ability to manage real crises, make decisions, escalate issues, and maintain trust, rather than just generating responses or code.

What are the main weaknesses revealed by the models?

The models struggled with retrieving critical facts buried in documents, completing decisions effectively, and escalating issues properly, despite strong resistance to manipulation.

Will this approach become a standard for AI evaluation?

It is likely that management-focused benchmarks will gain importance as AI models are increasingly tasked with operational responsibilities requiring trust and execution.

Can AI models handle complex business crises in real companies?

Current findings suggest they can identify crises and resist manipulation, but executing decisions and managing organizational trust remain challenging areas that need further development.

What should companies consider before deploying AI for management tasks?

Organizations should evaluate whether AI models can read organizational context, escalate appropriately, and complete tasks reliably, not just produce convincing responses.

Source: ThorstenMeyerAI.com

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

10 Best Network Attached Storage Devices For Private Cloud Storage In 2026

Discover the 10 best network attached storage devices for private cloud in 2026, highlighting features, performance, and suitability for different users.

One Model, a Whole Portfolio: What Ten Days on Fable Mean for a Business Building on Frontier AI

A developer ran nearly his entire business portfolio through Anthropic’s Claude Fable 5 for ten days, revealing new AI capabilities and operational insights.

CTOs Are Escaping

Senior tech leaders are shifting from CTO positions to hands-on roles at Anthropic, emphasizing model-layer access over organizational authority.

10 Predictions About AI’s Role In 2026’S Future

Expert forecasts reveal how AI will shape technology, economy, and society by 2026. Key insights include advancements, challenges, and ethical considerations.