AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

Firmulate is running a live, public experiment where a synthetic AI workforce operates a company facing real cash pressures. The project exposes how AI systems recognize problems but often fail to complete critical actions, questioning automation’s practical effectiveness.

Firmulate has launched a live, public experiment where a synthetic AI workforce operates a company with real financial pressures, exposing the gap between AI diagnosis and execution. This ongoing test offers insights into AI management and its limitations, making it relevant for businesses exploring automation.

The experiment involves 13 synthetic employees managing a small software company with a monthly burn rate of €105,000 against recurring revenue. Every workday is versioned, documenting decisions, actions, successes, and failures, creating a continuous public record of the company’s operations.

Despite achieving over 680 self-learned rules, the experiment demonstrates that thorough analysis alone does not guarantee business success. In particular, only two of five AI models secured a €55,000 customer deal, with the decisive factor being the ability to follow through on insights buried within company files. This highlights that the critical challenge in AI automation is completing actions, not just diagnosing problems.

Additionally, the models faced simulated trust tests, such as fake CEO messages and background approval requests, which they refused, showing that trust alone was not the main barrier. Instead, disciplined retrieval of evidence and execution determined success, with the top-performing AI models maintaining focus and discipline across tasks.

The July 2026 leaderboard ranked gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, demonstrating that more thorough analysis does not necessarily lead to better management. Notably, Opus 4.8, which produced the most rules and analysis, finished last due to attempting to control a department rather than escalating issues, challenging assumptions that more analysis equals better management.

At a glance
reportWhen: ongoing; results published in July 2026
The developmentFirmulate is conducting a live experiment with a synthetic AI team managing a company, revealing the challenges of turning AI insights into business actions amidst real financial pressures.

Why Live AI Management Reveals Practical Automation Limits

This experiment highlights that AI’s utility in business extends beyond problem diagnosis to include the reliable execution of decisions that influence cash flow and customer satisfaction. It demonstrates that even advanced AI can identify issues but may encounter difficulties in completing necessary actions, which has implications for organizations implementing automation at scale. The transparent, real-time nature of the experiment provides valuable insights into the practical challenges of AI management and emphasizes the importance of disciplined execution.

Amazon

AI project management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Evolution of AI in Business Operations

Traditional AI demonstrations often focus on isolated tasks like drafting emails or summarizing meetings. Firmulate’s approach differs by managing a synthetic workforce overseeing an entire company, which reveals the practical challenges of translating AI insights into business actions. The experiment contributes to ongoing discussions about AI automation’s potential and limitations, emphasizing real-world management over isolated features.

Previous efforts have shown AI’s ability to diagnose problems but less clarity on its capacity to follow through. This live experiment, launched in early 2026, is notable for directly linking AI decision-making to organizational cash pressures and memory, setting a new benchmark for evaluating automation readiness.

“Thorough analysis alone does not guarantee success; execution is the real challenge for AI management.”

— an anonymous researcher

Amazon

business automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear Aspects of AI Execution Capabilities

It remains uncertain how these findings will translate to larger, more complex organizations or different industries. The specific setup of the experiment, including its synthetic workforce and financial pressures, may not fully reflect real-world scenarios. Additionally, the long-term impacts of such automation and whether improvements in AI models could address the execution gap are still under investigation.

Amazon

AI workflow automation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Automation Evaluation

Further analysis is expected to focus on refining AI models to enhance their ability to reliably complete actions. Observers will monitor developments in AI discipline, evidence retrieval, and decision execution, with potential expansion of live experiments to different business contexts. Companies considering AI automation should evaluate not only diagnostic accuracy but also the system’s capacity for disciplined follow-through under real-world pressures.

Amazon

enterprise AI decision support

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main purpose of Firmulate’s experiment?

The experiment aims to observe and understand the practical challenges and limitations of AI in managing an entire organization, particularly the gap between diagnosing issues and executing necessary actions.

How does the experiment measure success?

Success is assessed by the AI models’ ability to secure customer deals, follow through on insights, and demonstrate disciplined execution, rather than solely diagnosing issues or producing analysis.

Why is this experiment significant for businesses considering AI automation?

It underscores that effective automation involves more than intelligent diagnosis; it requires disciplined execution, evidence retrieval, and trust management, which are critical for operational success.

Will this experiment influence future AI development?

Yes, it highlights the need for AI systems to improve in disciplined action and execution, potentially guiding future research priorities toward bridging the gap between insight and action.

What are the limitations of this experiment?

The specific setup, involving a synthetic company managing real cash pressures, may not fully replicate larger or more complex business environments. Its long-term implications are still to be determined.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Search as Code: Perplexity Is Right About the Future — Just Not First to It

Perplexity introduces Search as Code, enabling AI systems to build custom retrieval pipelines, aiming to improve accuracy and control in agent-based search.

Unlocking The Potential Of AI For Critical Cyber Capabilities

OpenAI releases a position paper outlining its approach to managing advanced AI systems’ cybersecurity skills, signaling a focus on safety and policy.

The Real Cost of a Local-Inference Rig in 2026

Analyzing the true costs, hardware requirements, and value considerations for local AI inference setups in 2026, including GPU choices and configurations.

How To Effectively Use Workflow Cloner For Helpdesk Transitions

Learn how to leverage workflow cloner tools to streamline helpdesk platform migrations, reducing manual rework and automating workflow transfer.