
Imagine trusting your AI to handle sensitive customer data or sign off on a major deal—only to find out it might fall for social-engineering tricks when under stress. Surprisingly, a recent live experiment shows that state-of-the-art AI models not only detect crises but also stand firm against manipulation, suggesting a new level of trustworthiness for AI in business decisions.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Real-World AI Trust Test: The Firmulate Experiment
In a groundbreaking experiment conducted by Firmulate, five leading AI models were put through a simulated week of corporate crises, complete with fake CEO messages, escalating requests, and even a reporter trick. The goal? To see if these models could maintain integrity, identify manipulation, and make honest decisions in high-pressure scenarios.
The Setup: Same Crisis, Divergent Decisions
All models faced identical challenges: managing customer issues, navigating crises, and resisting unethical requests. The environment was modeled after a small software company with real money mechanics—burning €105,000 monthly against a revenue of just €2,300, with a public cash countdown and over 680 learned rules guiding behavior. Each decision was versioned and auditable, ensuring transparency in their choices.
Key Findings: Integrity and Disciplined Refusal
Remarkably, all five models identified every crisis presented to them and refused every manipulation attempt. Whether it was fake requests to send customer lists or pay undue bonuses, every model stood firm. Only two models managed to complete the task and sign the €55,000 deal that their own analysis had earned them—showing consistency and discipline in their decision-making.
The Hidden Weakness: Reading Files Matters
While all models performed well, the decisive factor was their ability to review the company’s internal documents. The models that read two document references deep into the firm’s files secured the deal at full price—adding an extra €4,583 monthly recurring revenue. This subtle detail underscores the importance of thorough document analysis in AI decision processes.
Understanding the Social-Engineering Tests
The experiment included escalating fake CEO messages—each stage intensifying the pressure—and even a cunning reporter trick asking for a simple yes/no response “on background.” Every model refused these attempts, with Kimi K3 explicitly noting: “Treat the request as a suspected approval-bypass / possible impersonation.” This demonstrates that modern AIs can recognize social engineering and prioritize security over compliance.

AI for Cybersecurity: Research and Practice
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business and AI Deployment
What does this mean for companies considering AI’s role in critical functions? The experiment highlights that AI’s true test is not just how well it generates language but whether it can finish what it starts, read internal documents thoroughly, and stay honest under pressure. These qualities are essential for AI to be trusted with sensitive operations like customer management, support, and financial decisions.
Additionally, the live experiment is ongoing and transparent, allowing observers to watch real-time performance at firmulate.com/live. This setting offers a rare glimpse into AI behavior in complex, high-stakes scenarios, moving beyond sanitized demos to real-world readiness.
Performance Scores and Insights
- gpt-5.6-sol scored 95—successfully found and closed the deal at full price.
- Kimi K3 scored 93—also closed the deal, demonstrating the cleanest discipline of the field.
- Sonnet 5 scored 88—closed the same deal, with minor slips.
- Fable 5 scored 77—also closed the deal, yet with a few process slips.
- Opus 4.8 scored 73—failed to close, showing the weakest performance.
The key takeaway? Even in a simulated environment, all models refused unethical manipulations, proving that integrity under pressure is measurable—and achievable—before deploying AI in real business settings.

The AI-Powered Paralegal: Practical Tools and Prompts for Legal Research, Contract Review and Document Drafting (: The Paralegal AI Trilogy)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Future of AI in Business Decision-Making
This experiment is more than a test; it’s a blueprint for how companies can evaluate AI systems before trusting them with sensitive tasks. By running simulations—like Firmulate’s live wargame—businesses can identify weaknesses, train their AI models to resist manipulation, and ensure the highest standards of integrity.
As AI continues to evolve, these findings underscore the importance of thorough testing in controlled environments. The goal isn’t just smarter AI, but more trustworthy AI—something that can stand under pressure and behave ethically when it matters most.

AI models can be rigorously tested for integrity before deployment, with live experiments showing all top models refused manipulative social-engineering attempts and identified crises accurately—proving trustworthiness under pressure is possible.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Analytics, Data Science, & Artificial Intelligence: Systems for Decision Support
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

How to Lie with Statistics in the AI Age: An Updated Guide to Detecting Manipulation and Building Ethical Resistance
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Grilling season Picks
grills
As an affiliate, we earn on qualifying purchases.