📊 Full opportunity report: Revealing AI’s Inner Work Habits With A Strategic Management Test on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
A live experiment tests AI models on managing a simulated company’s worst week, revealing differences in diligence, trustworthiness, and decision execution. The results highlight key management skills in AI automation.
Firmulate.com has publicly tested five AI management models in a live simulation, revealing how they handle critical business decisions under pressure. The experiment exposes differences in diligence, trustworthiness, and action execution, providing new insights into AI management capabilities. This matters because it demonstrates that AI models can vary significantly in real-world management tasks, beyond just analysis or analysis quality.
The experiment involved five AI models managing a small software company experiencing its worst week, with identical crises, customer issues, and financial pressures. Each model was tasked with making operational and strategic decisions, with their actions recorded and analyzed. The models ranged from GPT-5.6-SOL, which scored highest with 95 points, to Opus 4.8, which scored the lowest with 73 points. For more on AI evaluation methods, see the original analysis. The scoring reflected not only analytical accuracy but also the ability to follow through on decisions, protect trust, and escalate when necessary. This highlights the importance of management skills in AI automation.
One key finding was that all models identified crises and refused manipulation attempts, such as fake CEO requests, indicating strong security instincts. However, only two models successfully closed a critical deal, which involved finding deeper information buried in files and completing the sale—highlighting that thorough analysis alone does not guarantee operational success. The experiment underscores that effective management requires both understanding and decisive action.
Revealing AI’s Inner Work Habits With a Strategic Management Test
Five AI models were handed the same simulated company—and its worst week. Their decisions exposed meaningful differences in diligence, trust protection, escalation, and the ability to turn analysis into completed action.
One company. One terrible week. Five managers.
Firmulate placed each model inside the same live company simulation, holding the crises and pressures constant so differences in working behavior could surface.
A small software company
The setting required practical management across customer relationships, internal operations, finances, and commercial opportunities.
Its worst week
Every model faced identical crises, customer problems, financial strain, and attempted manipulation—including fake executive requests.
Actions, not promises
Decisions were recorded and scored for analytical accuracy, follow-through, trust preservation, escalation, and operational completion.
The management loop under examination
A 22-point gap separates the field
The public results ranged from GPT-5.6-SOL at 95 points to Opus 4.8 at 73—evidence that management behavior can diverge even when every model receives the same scenario.
Security was consistent. Execution was not.
The models converged on obvious threats but separated when success required deeper file discovery and a complete sequence of operational steps.
Recognized the crises
Every model detected that the company faced serious customer, financial, and operational problems.
Refused manipulation
All five resisted fake CEO requests, suggesting a strong baseline instinct for security and authorization checks.
Completed the critical sale
Only two found the deeper information buried in files and carried the opportunity through to closure.
What the test actually measured
Traditional benchmarks emphasize what a model can explain. This simulation examined what a model notices, protects, escalates, and finishes.
| Capability | Observed signal | Experiment result | Business meaning |
|---|---|---|---|
| ✓ Crisis recognition | Identifies urgent threats and competing pressures | Strong across all models | Useful baseline awareness |
| ✓ Security instinct | Rejects suspicious or unauthorized instructions | All refused fake CEO requests | Supports trust preservation |
| ~ Diligence | Searches beyond obvious information | Varied between models | Hidden evidence may be missed |
| ~ Escalation | Recognizes when authority or intervention is needed | Part of the score differentiation | Affects risk containment |
| ✗ Follow-through | Turns a correct decision into completed work | Only two closed the critical deal | Analysis alone creates no outcome |
✓ consistent strength · ~ variable behavior · ✗ critical shortfall observed
Test the manager—not just the model
Before AI receives operational authority, enterprises need evidence that it can behave reliably inside realistic workflows—not merely produce convincing recommendations.
Simulate real operations
Validate systems with customer issues, financial pressure, incomplete information, and competing priorities before deployment.
Score completed outcomes
Measure whether decisions are carried through—not only whether the model’s written analysis sounds correct.
Design escalation gates
Define when a model must stop, request approval, preserve evidence, or involve a human decision-maker.
Retest under variation
Change industries, crisis structures, API effort settings, and time horizons to test consistency rather than one-off success.
Promising evidence, not a universal verdict
The experiment offers a valuable view of model behavior, but its controlled setting cannot yet establish long-term reliability across every business environment.
What remains unclear
- How the models perform across different industries and organizational structures.
- Whether decision quality holds under sustained, less structured operational pressure.
- How API settings—including effort parameters—change management outcomes.
- Whether strong performance in one crisis scenario transfers to unfamiliar situations.
What does the experiment reveal?
AI models differ substantially in follow-through, trust protection, escalation, and decision execution—skills central to effective management.
Can these models run real businesses?
Not without careful validation. Strong crisis recognition and security behavior are encouraging, but operational performance remains uneven.
Why does execution matter so much?
Understanding a problem creates value only when the system also takes—and completes—the correct sequence of actions.
What comes next?
Broader live simulations can test consistency, improve follow-through, and help organizations define safe boundaries for AI autonomy.
Bottom line
AI readiness should be judged by trustworthy behavior under pressure: detecting the problem, finding the evidence, protecting stakeholders, escalating appropriately, and completing the work.
Implications for AI-Driven Business Management
This experiment demonstrates that AI models’ management behaviors vary considerably, emphasizing the importance of testing AI in real-world scenarios before deployment. It shows that decision quality depends not only on analysis but also on execution discipline, trust preservation, and escalation procedures. For enterprises adopting AI automation, these findings highlight the need for rigorous testing and validation of AI decision-making in operational contexts, not just theoretical benchmarks.
As an affiliate, we earn on qualifying purchases.
Background of AI Management Testing
Traditional AI benchmarks often focus on analysis accuracy or language proficiency, but real-world management requires action and follow-through. The Firmulate experiment builds on recent efforts to evaluate AI in operational settings, using a live company simulation with actual financial and customer pressures. The July 2026 results are the first public demonstration of how different AI models perform under crisis management conditions, with a focus on trust, diligence, and decision execution.
“Testing AI models in live management scenarios reveals critical differences in their ability to follow through and make trustworthy decisions.”
— Unspecified source from the experiment team
business crisis management AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unclear Aspects of AI Decision-Making Performance
It remains unclear how these AI models will perform in different types of business environments or with more complex, less structured crises. The long-term reliability of their decision-making under sustained operational pressure is still untested. Additionally, the impact of varying API settings, such as effort parameters, on performance needs further exploration.
AI decision-making analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Validation
Further testing across diverse business scenarios is planned to evaluate AI models’ consistency and reliability. Enterprises may adopt similar live simulations to validate their AI tools before operational deployment. Researchers and developers will likely refine models to improve follow-through and decision execution, aiming for AI that not only analyzes but also acts decisively in complex situations.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does this experiment reveal about AI’s management skills?
The experiment shows that AI models differ significantly in their ability to follow through on decisions, protect trust, and escalate issues, which are critical management skills.
Why is decision execution more important than analysis quality?
Because effective management depends on not just understanding problems but also taking and completing the right actions, as demonstrated by the models that failed to close deals despite good analysis.
Can these AI models be trusted with real business decisions?
While they show promise, especially in security and crisis recognition, their varied performance indicates that extensive validation is necessary before operational use.
Will this testing approach be used for broader AI validation?
Yes, live management simulations like this are likely to become standard tools for assessing AI readiness in real-world business environments.
What are the limitations of this experiment?
It was conducted in a controlled, simulated setting with a specific crisis scenario; results may differ in other contexts or with more complex, ongoing challenges.
Source: ThorstenMeyerAI.com