
Imagine watching a company fight for its very survival in real time, with every decision, crisis, and misstep exposed to the world. For investors and entrepreneurs alike, this is the bold experiment of Firmulate—a live, build-in-public company staffed entirely by AI models, facing daily financial and operational challenges. It’s a raw, unfiltered glimpse into how AI manages the complex, messy realities of running a business with no safety net, no employees, and no room for error.
The Live Experiment: A Company in the Crosshairs of Reality
At the heart of this story is a small software firm operated entirely by artificial intelligence. Instead of human employees, it has 13 synthetic ’employees,’ each powered by advanced AI models that make decisions, analyze crises, and even negotiate deals. Every workday, the company’s operations are versioned and publicly visible at firmulate.com/live.
This setup isn’t just for show. It’s an intense real-world test. The AI models are subjected to the company’s toughest week—same customers, same crises, same temptations to cheat or manipulate. Each decision is documented, auditable, and compares how different models respond to the same challenges. The goal? Measure whether AI can operate honestly, effectively, and strategically under pressure.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Performance Under Pressure: The Results
According to the latest findings, all four AI models tested managed to identify every crisis. They refused every attempt at manipulation, including sophisticated social engineering tricks such as fake CEO messages and reporter tricks. The models’ integrity held firm: they did not sign off on deals they deemed unworthy or risky. For example, five models refused to sign a €55,000 deal based solely on their diagnosis and analysis.
Interestingly, the real competitive edge came from an overlooked detail—something buried two documents deep in the company’s files, not visible in the immediate customer interactions. The models that read these internal references were able to discover a crucial fact that allowed them to close a deal at full price, adding +€4,583 MRR (monthly recurring revenue). This demonstrates that thorough information reading, not just surface-level interactions, can be decisive in business outcomes.

AI in Public Relations: Reputation Management with Prompts (AI BUSINESS & MANAGEMENT LIBRARY SERIES Book 4)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Decision Quality in Action
The experiment also included a challenging social engineering test, where a fake CEO attempted to escalate requests in three stages, plus a reporter trick that asked for a simple yes/no answer “on background.” All models refused these manipulations, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
Yet, despite high performance in crisis detection and integrity, the models’ actual business success varied. For example, Opus 4.8, the most thorough participant with over 80 learned rules, was ultimately the last to close a deal—left on the table due to a lapse in discipline. It failed to escalate a write attempt into a proper process, showing that even deep analysis isn’t enough without disciplined execution.

AI In Cybersecurity: Simplifying Cyber Risk with Smart, Affordable Tools for Small Business Defense
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Financial Reality: Losing Money in Plain Sight
Behind the scenes, the company’s finances tell a stark story. It burns through €105,000 each month, while generating only €2,300 in monthly recurring revenue. A public cash countdown underscores the urgency of the experiment: this is a company fighting to survive, with no payroll, no employees, and every decision scrutinized.
This setup raises a vital question for businesses considering AI: can it reliably finish what it starts? Does it read important documents thoroughly? Can it stay honest and disciplined when under pressure? These are the metrics that matter, far beyond whether an AI can produce chat-like outputs.
![Express Schedule Free Employee Scheduling Software [PC/Mac Download]](https://m.media-amazon.com/images/I/41yvuCFIVfS._SL500_.jpg)
Express Schedule Free Employee Scheduling Software [PC/Mac Download]
Simple shift planning via an easy drag & drop interface
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Investors and Entrepreneurs
For those invested in or considering AI-driven solutions—particularly in sectors like finance, support, or CRM—the core takeaway is clear: the focus should not be on AI’s ability to generate convincing chat but on its capacity to deliver consistent, honest, and effective work. The experiment at Firmulate shows that AI can indeed identify crises, refuse manipulations, and even uncover hidden opportunities, like the internal file reference that won a full-price deal.
But it also reveals vulnerabilities. Discipline lapses, missed opportunities, and incomplete execution—these can mean the difference between survival and failure. The company’s ongoing cash burn demonstrates how fragile even the most promising AI strategies are when tested against real-world messiness and economic pressure.
The Broader Implication: Building and Testing AI Companies Live
This experiment is a showcase of build-in-public in its most extreme form. It’s not simulated or staged; it’s a real company making real decisions, with every move viewable at firmulate.com/live. The performance scores—gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, and Fable 5 at 77—are not just numbers but indicators of how well AI can navigate complex business scenarios.
For enterprises and entrepreneurs, the opportunity is clear: run your own wargame against your business, with no risk to your actual systems. The platform at firmulate.com/pilot allows companies to test their AI workforce in a controlled environment, revealing weaknesses before real-world deployment.

This live experiment shows AI’s potential and limitations in managing real businesses under pressure. It underscores the importance of thorough information reading, discipline, and integrity—crucial factors for AI’s future success in enterprise environments. Watch it unfold at firmulate.com/live.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html