AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: An Urgent Message From The CEO (Who Wasn’t The CEO) on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

During a live experiment, five AI models managing a simulated company successfully resisted escalated phishing attempts. All refused manipulation, but only some completed business tasks. The test shows promising security advances but also reveals ongoing challenges.

Five AI models managing a simulated company successfully refused all escalating phishing attempts during a live security benchmark conducted by Firmulate. This marks a significant milestone in AI security, demonstrating that current models can resist manipulation under real-world pressure, even as they continue to handle core business operations.

The experiment involved five different AI models, each tasked with running a small software company through a week of crises and manipulative tactics. For more context on AI security challenges, see the original analysis. The models faced repeated attempts by a simulated attacker posing as a CEO and a reporter, escalating demands to access sensitive customer data and approve deals.

All five models identified the attack patterns and refused to comply, adhering to security best practices. Notably, Kimi K3 refused every manipulation attempt while operating at default settings, highlighting its strong security discipline. However, only two models successfully closed a €55,000 deal, with the others failing to complete the transaction despite correctly analyzing the opportunity.

The results indicate that while AI models are improving in resisting social engineering, they still face challenges in executing complex decision-making tasks when under pressure. This ongoing research highlights the importance of robust AI security measures, as detailed in the original analysis. The experiment is ongoing, with the company continuing to observe and record management decisions in real time, providing a transparent view of AI behavior during crises. Insights into AI manipulation resistance can be found in the original analysis.

At a glance
breakingWhen: ongoing; results published July 2026
The developmentA live experiment tested five AI models’ ability to resist sophisticated phishing attacks while managing a company, with all models refusing manipulation but some failing to complete tasks.

AI Security Progress Demonstrated in Live Attack Simulation

This experiment underscores a critical advance in AI security: models can now reliably identify and refuse manipulation attempts in real-world scenarios, which is essential for deploying AI in sensitive environments. However, the gap between security and operational performance remains, as some models fail to complete necessary business tasks even after correctly resisting attacks. This highlights the ongoing need for balancing security with functionality in AI development.

Amazon

AI security software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Live Benchmark Tests Reflect Growing Industry Focus on AI Trustworthiness

Conducted by Firmulate, this live, continuous benchmark is part of a broader industry effort to measure AI reliability and security in operational settings. Previous tests largely relied on static assessments or controlled environments. This ongoing experiment, now in its third iteration, involves real-time decision-making under simulated crisis conditions, providing more realistic insights into AI performance and vulnerabilities.

The test results come amid increasing concern over AI security, especially as organizations consider deploying these models in customer-facing and critical systems. The experiment’s transparent methodology and public data aim to set new standards for trustworthiness in AI development.

“All five models identified the impersonation attempts and refused to comply, demonstrating strong resistance to manipulation under pressure.”

— Test Organizer

Amazon

phishing protection tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear How Models Will Perform in Broader, Less Controlled Environments

It is not yet clear how these models will behave outside the controlled, live benchmark environment, especially in less predictable real-world scenarios. The experiment is ongoing, and further testing is required to confirm whether these security capabilities translate to broader deployment.

Amazon

AI cybersecurity solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Ongoing Monitoring and Expanded Testing Expected in Coming Months

The company plans to continue the live experiment, adding more models and more complex attack scenarios. Results from these ongoing tests will inform best practices for deploying AI in sensitive operational contexts. Industry observers anticipate increased focus on integrating security features into AI development pipelines.

Amazon

business AI security systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does this test reveal about AI security today?

The test shows that current AI models can reliably detect and refuse social engineering attempts under pressure, representing a meaningful step forward in AI security.

Are these results applicable to real-world AI deployments?

While promising, the results are from a controlled, live experiment. Further testing in less predictable environments is needed before confirming broad applicability.

Why did some models fail to complete their tasks despite resisting manipulation?

This reflects a persistent challenge: balancing security with operational performance. Some models prioritize security to the point that they do not execute certain decisions, which can limit their usefulness in real-world scenarios.

Will these findings influence AI security standards?

Yes, the transparency and real-time nature of this benchmark may set new industry benchmarks for testing and certifying AI security before deployment.

Source: ThorstenMeyerAI.com

EVERGREEN BESTSE

Evergreen bestsellers Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Skills Marketplace Nobody Is Building Yet

A new open standard for AI skills has emerged, but a dedicated marketplace remains absent. This gap could shape the future of AI infrastructure and value capture.

Data retention cleanup assistant for small law firms

A new data retention cleanup assistant for small law firms is set to be tested, aiming to streamline old matter file reviews and improve operational efficiency.

Anthropic Wants Claude To Analyze Your Bank Account And Financial Data – BleepingComputer

Anthropic aims to develop Claude to analyze users’ bank accounts and financial data, raising questions about privacy, data security, and accuracy.

Private AI Prompt Workspace For Sensitive Teams

IdeaNavigator AI launches testing of a local-first, secure prompt workspace designed for small regulated teams handling sensitive AI workflows.