
Imagine a world where your AI system doesn’t just chat nicely, but actually runs a company — making tough decisions, resisting manipulation, and closing real deals. Sounds futuristic? Well, it’s happening now. At Firmulate, AI models are put through a rigorous live experiment, acting as virtual CEOs during their worst week. The question isn’t just if they can talk—they’re being tested on whether they can truly manage, stay honest, and deliver results like real human managers.
Open a free Amazon Business account
Business pricing, bulk buying and tax-exempt orders.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: Putting AI to the Test in Real Business Conditions
Every AI model in the experiment was tasked with running a small software company facing the kind of crises and temptations that would make any manager sweat. From customer complaints to internal manipulation attempts, each model faced identical challenges, with every decision carefully recorded and made transparent for analysis. This setup isn’t a mere demo; it’s a full-blown simulation of real-world business management, happening live at firmulate.com.
Who Were the Competitors?
- GPT-5.6-sol scored highest at 95 points, identifying hidden information that sealed a €55,000 deal.
- Kimi K3, the newcomer, scored 93, also closing the deal with impeccable discipline.
- Sonnet 5 and Fable 5 scored 88 and 77 respectively, both closing the deal but with noticeable process slips.
Interestingly, all models resisted manipulation attempts, such as staged CEO messages and media inquiries, refusing to be tricked into unethical acts. For instance, all five models rejected a staged fake CEO message escalating tensions, reasoning that it was suspicious or impersonation.
The Hidden Factor: Reading Between the Lines
A critical insight emerged: the decisive advantage went beyond surface-level decisions. It was the models’ ability to read deep into company files, uncover buried facts, and act on them. Those who identified two document references deep within internal files won the full-price deal, worth an additional €4,583 in monthly recurring revenue.
What About the Real Business?
Firmulate’s experiment runs a live, small software company with 13 synthetic employees, managing real money mechanics—burning €105,000 a month against only €2,300 MRR. The company works every business day, with over 680 self-learned rules and every decision versioned and auditable. This isn’t a simulation; it’s an operational company under AI management, open for viewers to watch at firmulate.com/live.
AI decision-making software for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What the Results Tell Us About AI Management Personalities
The experiment reveals that AI models have measurable management characters. For example, Opus 4.8, the most thorough participant with over 80 rules learned, analyzed every option deeply but ultimately left a critical close untouched—showing discipline but also hesitation. Meanwhile, the Kimi K3 model ran without an effort parameter, which made it more disciplined and decisive, closing deals at a high rate. Less disciplined models slipped in process slips but still managed to close, highlighting differences in operational style rather than ability.
Why Does This Matter?
For businesses considering AI-driven decision-making, the question isn’t just whether the AI can generate human-like dialogue. Instead, it’s whether the AI can finish what it starts, read critical information, and stay honest under pressure. The models’ ability to resist manipulation and uncover hidden truths—like the buried document references—could determine whether AI manages your CRM, support queue, or forecasts reliably.
As an affiliate, we earn on qualifying purchases.
Share the Power of Wargaming Your AI Workforce
Through this live experiment, firms can run their own ‘wargames’ against a read-only export of their business—testing how AI would perform in their actual environment without affecting real systems. It’s a cost-effective way to see if your AI can be trusted to deliver results that matter, not just produce convincing chat.
The Final Word
The experiment at Firmulate shows that AI models can be more than just chatbots—they can actively manage, make tough decisions, and stay honest amid crises. The key is understanding their management style, discipline, and ability to read deep into your business data. As AI takes a bigger role in enterprise workflows, the question is: which model will lead your team to success?

The live experiment at Firmulate proves that AI models can manage real business crises, stay honest under manipulation, and close deals—values crucial for enterprise trust. Knowing a model’s decision style helps businesses pick the right AI for their needs.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI-powered business crisis management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Back to school Picks
back to school
As an affiliate, we earn on qualifying purchases.