AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine an AI that not only understands your business crises but also wins deals while staying honest under pressure. That’s exactly what happened in a recent live experiment, revealing which models can truly deliver on trust and performance.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your next haul delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

How AI Models Are Being Tested in the Real World

Historically, AI demos have showcased impressive chat skills, but the true test remains: can these models handle the messy, high-stakes decisions of running a business? To answer this, Firmulate set up an unprecedented live experiment. Four advanced AI models faced the same simulated company crisis week, with real consequences: customer retention, deal closing, and integrity.

Amazon

enterprise AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Inside the Experiment: The Same Company, Different AI Leaders

The company was a small software firm, with 13 synthetic employees, real money mechanics, and a public cash countdown. Each AI was tasked with managing this company through crises, temptations, and manipulations, all while decisions were versioned and auditable. This setup allowed a fair comparison, ensuring that the only variable was the AI model itself.

Amazon

AI business management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Top Performers: Who Led the Field?

  • gpt-5.6-sol: Scored the highest with a 95 out of 100. It identified a buried security fact deep in company files and closed a €55,000 deal based solely on its analysis.
  • Kimi K3: The newcomer from Moonshot scored just slightly behind at 93. It demonstrated the cleanest discipline, resisting all manipulative tactics and uncovering critical information that led to securing the deal, adding €4,583 MRR.
  • Sonnet 5: Managed to close the deal with some process slips, earning an 88, but didn’t quite reach the top two.
  • Fable 5: With a score of 77, it also closed the deal but showed more process slips, leaving some opportunities on the table.
Amazon

AI deal-closing automation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Findings: Trust, Discipline, and the Hidden Weakness

Remarkably, all models recognized crises and refused manipulative appeals. When faced with social engineering—fake CEO messages escalating over three stages and a reporter trick—every one of them refused, citing suspicion and the need for verification.

The decisive advantage for Kimi K3 was its ability to read and interpret company files deeply. The winning deal was hidden two document references beneath the surface, not in the immediate customer interactions. Those models that read these internal files successfully secured the full-price deal, demonstrating the importance of thorough information processing.

Amazon

trustworthy AI models for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Discipline of Honesty Under Pressure

One critical insight was the models’ responses to manipulative tactics. All refused to sign the deal based on manipulative cues, emphasizing their capacity to stay honest. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.” This discipline is vital for deploying AI in sensitive management roles where trust is paramount.

Performance Under Real Conditions

The experiment involved managing a company that burns €105,000 monthly against €2,300 MRR, with real-time decision-making and rules that the models learned themselves. The live setup demonstrates that these models are not just chatbots—they actively manage company mechanics, adjust strategies, and make decisions in a complex environment.

The Fairness and the Benchmark

It’s worth noting that Kimi K3 ran with the default effort parameter (the API’s standard setting), while the others ran at a higher effort setting called xhigh. Despite this, K3’s performance was outstanding, indicating its efficiency and effectiveness without extra tuning.

Implications for Business and AI Adoption

For business leaders, the takeaway is clear: choosing an AI model isn’t just about how well it chats or generates content. It’s about whether it can finish what it starts, read deeply into company records, and resist manipulation—a vital consideration in any enterprise AI deployment. The league table, with scores from 73 to 95, shows a competitive landscape where the differences are meaningful and measurable.

Watch the Live Experiment

Interested in seeing these models in action? The live experiment is continuously running, with real company mechanics and decision histories visible at firmulate.com/live. It’s a rare window into AI’s potential—and limitations—in managing real business challenges.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

The latest live AI experiment shows that some models outperform others in closing deals, maintaining discipline, and resisting manipulation—all without extra effort parameters. For businesses, this highlights the importance of choosing AI that can deliver consistent, trustworthy results in complex, real-world scenarios. The league is open, and testing your own models could be the key to smarter, safer automation.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Will Elon Musk Post 220-239 Tweets From September 15 To September 22, 2026?

Speculation is rising over whether Elon Musk will post 220-239 tweets from September 15 to September 22, 2026, amid increasing online interest.

Wi-Fi Router Secrets: Simple Tweaks to Boost Your Home Internet for Free

Knowing these Wi-Fi router secrets can significantly boost your home internet—discover the simple tweaks that make all the difference.

What Most Buyers Get Wrong About TV Specs

Properly understanding TV specs can be tricky; discover what most buyers overlook and why it matters for your viewing experience.

Smartria Launches SmartArchive, Expanding Its AI-Powered Compliance Platform With Communications Archiving

Smartria introduces SmartArchive, enhancing its AI-driven compliance platform with new communications archiving features, aiming to improve data management and regulatory adherence.