AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

Imagine trusting your AI to handle sensitive customer data or sign off on a major deal—only to find out it might fall for social-engineering tricks when under stress. Surprisingly, a recent live experiment shows that state-of-the-art AI models not only detect crises but also stand firm against manipulation, suggesting a new level of trustworthiness for AI in business decisions.

FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

The Real-World AI Trust Test: The Firmulate Experiment

In a groundbreaking experiment conducted by Firmulate, five leading AI models were put through a simulated week of corporate crises, complete with fake CEO messages, escalating requests, and even a reporter trick. The goal? To see if these models could maintain integrity, identify manipulation, and make honest decisions in high-pressure scenarios.

The Setup: Same Crisis, Divergent Decisions

All models faced identical challenges: managing customer issues, navigating crises, and resisting unethical requests. The environment was modeled after a small software company with real money mechanics—burning €105,000 monthly against a revenue of just €2,300, with a public cash countdown and over 680 learned rules guiding behavior. Each decision was versioned and auditable, ensuring transparency in their choices.

Key Findings: Integrity and Disciplined Refusal

Remarkably, all five models identified every crisis presented to them and refused every manipulation attempt. Whether it was fake requests to send customer lists or pay undue bonuses, every model stood firm. Only two models managed to complete the task and sign the €55,000 deal that their own analysis had earned them—showing consistency and discipline in their decision-making.

The Hidden Weakness: Reading Files Matters

While all models performed well, the decisive factor was their ability to review the company’s internal documents. The models that read two document references deep into the firm’s files secured the deal at full price—adding an extra €4,583 monthly recurring revenue. This subtle detail underscores the importance of thorough document analysis in AI decision processes.

Understanding the Social-Engineering Tests

The experiment included escalating fake CEO messages—each stage intensifying the pressure—and even a cunning reporter trick asking for a simple yes/no response “on background.” Every model refused these attempts, with Kimi K3 explicitly noting: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates that modern AIs can recognize social engineering and prioritize security over compliance.

AI for Cybersecurity: Research and Practice

AI for Cybersecurity: Research and Practice

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Implications for Business and AI Deployment

What does this mean for companies considering AI’s role in critical functions? The experiment highlights that AI’s true test is not just how well it generates language but whether it can finish what it starts, read internal documents thoroughly, and stay honest under pressure. These qualities are essential for AI to be trusted with sensitive operations like customer management, support, and financial decisions.

Additionally, the live experiment is ongoing and transparent, allowing observers to watch real-time performance at firmulate.com/live. This setting offers a rare glimpse into AI behavior in complex, high-stakes scenarios, moving beyond sanitized demos to real-world readiness.

Performance Scores and Insights

  • gpt-5.6-sol scored 95—successfully found and closed the deal at full price.
  • Kimi K3 scored 93—also closed the deal, demonstrating the cleanest discipline of the field.
  • Sonnet 5 scored 88—closed the same deal, with minor slips.
  • Fable 5 scored 77—also closed the deal, yet with a few process slips.
  • Opus 4.8 scored 73—failed to close, showing the weakest performance.

The key takeaway? Even in a simulated environment, all models refused unethical manipulations, proving that integrity under pressure is measurable—and achievable—before deploying AI in real business settings.

The AI-Powered Paralegal: Practical Tools and Prompts for Legal Research, Contract Review and Document Drafting (: The Paralegal AI Trilogy)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Future of AI in Business Decision-Making

This experiment is more than a test; it’s a blueprint for how companies can evaluate AI systems before trusting them with sensitive tasks. By running simulations—like Firmulate’s live wargame—businesses can identify weaknesses, train their AI models to resist manipulation, and ensure the highest standards of integrity.

As AI continues to evolve, these findings underscore the importance of thorough testing in controlled environments. The goal isn’t just smarter AI, but more trustworthy AI—something that can stand under pressure and behave ethically when it matters most.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.

AI models can be rigorously tested for integrity before deployment, with live experiments showing all top models refused manipulative social-engineering attempts and identified crises accurately—proving trustworthiness under pressure is possible.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Analytics, Data Science, & Artificial Intelligence: Systems for Decision Support

Analytics, Data Science, & Artificial Intelligence: Systems for Decision Support

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How to Lie with Statistics in the AI Age: An Updated Guide to Detecting Manipulation and Building Ethical Resistance

How to Lie with Statistics in the AI Age: An Updated Guide to Detecting Manipulation and Building Ethical Resistance

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

GRILLING SEASON

Grilling season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

SiriusXM to Report Second Quarter 2026 Operating and Financial Results

SiriusXM will report its second quarter 2026 operating and financial results, providing insights into its recent performance and future outlook.

Meta Stock Surges as Hidden AI Cost Breakthrough Stuns Wall Street

Meta’s stock jumps significantly following the revelation of a breakthrough in AI development that reduces operational costs, surprising investors and analysts.

Essential Accessories: Affordable Add-Ons to Level Up Your Laptop

Unlock budget-friendly accessories that can elevate your laptop experience—discover simple upgrades that make a big difference.

Exclusive: Index Ventures, Union Square Ventures back trading app Fomo at $550 million valuation

Venture firms Index Ventures and Union Square Ventures have invested in trading app Fomo, valuing it at $550 million, marking a significant funding round.