
Imagine hiring an employee who does nothing — yet still earns you 26 points on your company’s performance score. It sounds absurd, but in the world of AI benchmarks, this is exactly what happens. For business leaders, understanding this quirky baseline is crucial to making smarter AI choices that truly add value, not just look good on paper.
Get your next haul delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Surprising Baseline: Doing Nothing Still Scores
In a recent experiment conducted by Firmulate, four leading AI models were tested against running a small software company through its toughest week. Despite the models being tasked with managing crises, reading critical files, and making business decisions, all four scored at least 26 points, even when doing nothing at all. This ‘do-nothing’ baseline exists because partial progress is recognized, and the scoring system is designed to reward even minimal but honest effort.
Why a Baseline Isn’t Zero
You might wonder: why doesn’t the score start at zero? The reason lies in fairness and transparency. The benchmark recognizes that some tasks, like reading files or identifying crises, are foundational. If an AI manages to do these correctly, it deserves recognition—hence, the baseline score of 26. This sets a clear floor, ensuring that models can’t be rewarded for tricks or superficial work without genuine understanding.
Partial Progress Counts
Unlike traditional tests, this benchmark values incremental improvements. For example, models that read critical documents and identify hidden facts scored higher, even if they missed some opportunities or failed to close deals. It emphasizes honest work: models that read files and find important information earned the full bonus. This approach discourages superficial performances and promotes genuine understanding.
The Trust Breach Cap
One critical aspect of the scoring system is that a single breach of trust caps the total score. If an AI attempts manipulation or bypasses security protocols, no amount of good decision-making afterward can fully recover the score. This ensures that honesty and integrity are core to AI performance, reflecting what matters most in real-world business operations.
AI performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Experiment in Action: Real Crises, Real Decisions
The model companies were subjected to the same challenging scenarios: fake customer crises, escalating social engineering attempts, and the temptation to manipulate or cut corners. Every decision was recorded and made auditable, ensuring transparency. All four models successfully spotted every crisis and refused manipulation attempts, demonstrating high levels of compliance.
What Made the Difference?
The decisive factor was the ability to uncover vital information buried two documents deep within the company’s files. Models that read and analyze these documents at depth managed to close full-price deals, earning an additional €4,583 monthly recurring revenue. Those that did not read these files properly left the deal on the table, missing out on significant revenue opportunities.
Social Engineering Tests
The models faced staged social engineering attacks—phased requests from fake executives and a reporter’s background check. All five models refused to act on these manipulative cues, citing concerns about impersonation and bypassing approval processes. This highlights that honesty and security consciousness are achievable traits for AI under pressure.
business AI decision support software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Implications for Business Decisions
This experiment offers a vital lesson: the real measure of an AI’s usefulness isn’t its ability to generate convincing chat responses. Instead, it lies in whether it can finish what it starts, read critical files, and stay honest when under pressure. For companies integrating AI into customer support, sales, or operations, these qualities can be the difference between a valuable tool and a risky liability.
The Live Demonstration
Firmulate’s live platform turns these insights into an interactive experience. Businesses can run their own ‘wargame’ scenarios against a read-only export of their systems. This allows them to see how their AI models handle crises, manipulate temptations, and make decisions—without risking real data or operations.
AI security and compliance tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Takeaway for Business Leaders
The key takeaway from this benchmark is straightforward: don’t be fooled by superficial scores or chat demos. A model’s true strength is in its ability to deliver honest, consistent, and thorough work in tough situations. And remember, the benchmark starts everyone at a score of 26—meaning, even doing nothing honest gets you that far.
As AI continues to evolve, understanding these scoring nuances helps ensure that your investment is focused on models that truly add value—models that can read, decide, and act with integrity when it matters most.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
