
Imagine a company with no human employees, battling daily crises and losing thousands of euros every month, yet still managing to strike a €55,000 deal. Sounds impossible? Welcome to the world of Firmulate, a groundbreaking experiment in AI-driven business management that’s unfolding in real time, right before your eyes.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Living Experiment: A Company Run by AI Models
At the heart of this experiment is a small, real-world software company operated entirely by artificial intelligence. It’s a public, transparent venture where every decision, crisis handling, and negotiation is recorded and posted online. The company’s daily operations involve 13 synthetic employees making choices based on a set of 680+ self-learned rules, all while burning €105,000 per month against a modest monthly recurring revenue (MRR) of €2,300.
What makes this setup extraordinary is its transparency and the fact that it’s not just a simulation—this is a real, functioning company with a public cash countdown and detailed decision logs. Every workday, its progress is versioned and accessible for anyone to watch, analyze, or even challenge.

Building AI-Powered Products: The Essential Guide to AI and GenAI Product Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Testing AI Models in the Real Business Arena
Four leading AI models were tasked with running this company through its worst week—a week filled with customer crises, internal dilemmas, and manipulation attempts. These models had to read customer files, diagnose problems, pitch solutions, and even resist social engineering attempts such as fake CEO requests or reporter tricks.
Despite their different architectures and training parameters—ranging from the most thorough, Opus 4.8, to the more straightforward Kimi K3—all four models identified every crisis and refused every manipulation attempt. The twist? Only two of them managed to close the key €55,000 deal that their own analysis had earned, with full understanding of the company’s internal insights.

AI for Public Relations: A How-To Guide for Implementation and Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: The Buried Fact
Digging deeper into the company’s files, the models that read and understood the full context were the ones that won the deal at full price. The critical information was buried two document references deep within the company’s own files—yet only the models that accessed this buried knowledge succeeded in closing the deal at its true value, adding €4,583 in monthly recurring revenue.

As an affiliate, we earn on qualifying purchases.
Social Engineering and Ethical Resistance
The experiment also tested whether the AI would fall for social engineering tactics, like fake CEO messages or background-only approval requests. All five models involved in this test refused to comply, with Kimi K3 explicitly noting: “Treat the request as a suspected approval-bypass / possible impersonation.”

AI Phishing, Social Engineering & Fraud: How Criminals Use AI to Manipulate, Steal & Deceive (The AI Cybersecurity)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Reality of the Live Company
This isn’t just a demo or a simulation. The live company, accessible at firmulate.com/live.html, runs every business day. Its operations, decisions, and struggles are all transparent and viewable. Each decision is carefully versioned, providing a detailed audit trail that reveals not only how AI handles crises but also where it falters.
Performance and Lessons from the Models
The most detailed participant, Opus 4.8, with over 80 learned rules and deep analyses, was the last to close the deal—leaving some opportunities on the table due to discipline slips, like writing attempts into a locked department instead of escalating. Meanwhile, Kimi K3, which ran without an effort parameter (essentially default settings), performed almost as well in terms of fairness, even if slightly less disciplined.
The key takeaway is clear: AI models can reliably identify crises, resist manipulation, and even close deals—if they truly read the internal files and follow disciplined processes. Yet, the experiment exposes the gap between what AI can do in controlled demos and its performance in a real, money-losing environment fighting for survival.
Why You Should Care
If AI agents are going to integrate with your CRM, support channels, or forecasting tools, the question isn’t just about how well they generate text or handle conversations. It’s whether they can follow through, read your internal data, refuse manipulation, and deliver actual useful work that contributes to your bottom line.
The experiment’s real-world setup and transparent data serve as a stark reminder: Building trustworthy AI isn’t just about natural language prowess. It’s about their ability to execute, stay honest under pressure, and make decisions that matter—especially when the stakes are high and the company is losing money every day.

Watching a real, live company run entirely by AI models reveals vital insights: trustworthiness, discipline, and thorough internal data access are key to AI’s business potential. Despite financial struggles, the experiment proves that AI can win critical deals and resist manipulation—if it has the right data and discipline. For companies considering AI integration, the question is not just about chat quality, but about whether the AI can deliver consistent, honest results in real-world scenarios.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Baby shower & registry season Picks
baby registry must-haves
As an affiliate, we earn on qualifying purchases.