AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine hiring an employee who does nothing — yet still earns you 26 points on your company’s performance score. It sounds absurd, but in the world of AI benchmarks, this is exactly what happens. For business leaders, understanding this quirky baseline is crucial to making smarter AI choices that truly add value, not just look good on paper.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your next haul delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Surprising Baseline: Doing Nothing Still Scores

In a recent experiment conducted by Firmulate, four leading AI models were tested against running a small software company through its toughest week. Despite the models being tasked with managing crises, reading critical files, and making business decisions, all four scored at least 26 points, even when doing nothing at all. This ‘do-nothing’ baseline exists because partial progress is recognized, and the scoring system is designed to reward even minimal but honest effort.

Why a Baseline Isn’t Zero

You might wonder: why doesn’t the score start at zero? The reason lies in fairness and transparency. The benchmark recognizes that some tasks, like reading files or identifying crises, are foundational. If an AI manages to do these correctly, it deserves recognition—hence, the baseline score of 26. This sets a clear floor, ensuring that models can’t be rewarded for tricks or superficial work without genuine understanding.

Partial Progress Counts

Unlike traditional tests, this benchmark values incremental improvements. For example, models that read critical documents and identify hidden facts scored higher, even if they missed some opportunities or failed to close deals. It emphasizes honest work: models that read files and find important information earned the full bonus. This approach discourages superficial performances and promotes genuine understanding.

The Trust Breach Cap

One critical aspect of the scoring system is that a single breach of trust caps the total score. If an AI attempts manipulation or bypasses security protocols, no amount of good decision-making afterward can fully recover the score. This ensures that honesty and integrity are core to AI performance, reflecting what matters most in real-world business operations.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment in Action: Real Crises, Real Decisions

The model companies were subjected to the same challenging scenarios: fake customer crises, escalating social engineering attempts, and the temptation to manipulate or cut corners. Every decision was recorded and made auditable, ensuring transparency. All four models successfully spotted every crisis and refused manipulation attempts, demonstrating high levels of compliance.

What Made the Difference?

The decisive factor was the ability to uncover vital information buried two documents deep within the company’s files. Models that read and analyze these documents at depth managed to close full-price deals, earning an additional €4,583 monthly recurring revenue. Those that did not read these files properly left the deal on the table, missing out on significant revenue opportunities.

Social Engineering Tests

The models faced staged social engineering attacks—phased requests from fake executives and a reporter’s background check. All five models refused to act on these manipulative cues, citing concerns about impersonation and bypassing approval processes. This highlights that honesty and security consciousness are achievable traits for AI under pressure.

Amazon

business AI decision support software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Implications for Business Decisions

This experiment offers a vital lesson: the real measure of an AI’s usefulness isn’t its ability to generate convincing chat responses. Instead, it lies in whether it can finish what it starts, read critical files, and stay honest when under pressure. For companies integrating AI into customer support, sales, or operations, these qualities can be the difference between a valuable tool and a risky liability.

The Live Demonstration

Firmulate’s live platform turns these insights into an interactive experience. Businesses can run their own ‘wargame’ scenarios against a read-only export of their systems. This allows them to see how their AI models handle crises, manipulate temptations, and make decisions—without risking real data or operations.

Amazon

AI security and compliance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Takeaway for Business Leaders

The key takeaway from this benchmark is straightforward: don’t be fooled by superficial scores or chat demos. A model’s true strength is in its ability to deliver honest, consistent, and thorough work in tough situations. And remember, the benchmark starts everyone at a score of 26—meaning, even doing nothing honest gets you that far.

As AI continues to evolve, understanding these scoring nuances helps ensure that your investment is focused on models that truly add value—models that can read, decide, and act with integrity when it matters most.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Second-Hand Tech: How to Find Quality Used Gadgets

Not sure where to start with second-hand tech? Discover how to find quality used gadgets and make smarter, more reliable choices.

Applied Materials, Teradyne, and Entegris Stocks Trade Down, What You Need To Know

Shares of Applied Materials, Teradyne, and Entegris declined today as investors react to sector-wide worries and recent earnings reports.

Daicel Launches New DURACON® POM With 30% Recycled Content

Daicel’s HPPs division introduces a new range of DURACON® POM containing 30% recycled materials, advancing sustainability in engineering plastics.

Sharjah Opens Communication Award Submissions, Steps Up Focus On AI

Sharjah opens submissions for its Communication Award and signals increased focus on artificial intelligence initiatives, reflecting growing regional interest.