
Imagine a company with no employees, losing €105,000 each month, yet being run entirely by AI models under public scrutiny—welcome to one of the most radical experiments in business automation today.
The Living Experiment: A Business in Perpetual Crisis
At the heart of this bold venture is a live, publicly accessible simulation where 13 synthetic “employees” manage a small software company. This isn’t just a demo; it’s a real-time experiment with real money mechanics, a €105,000 monthly burn against €2,300 in monthly recurring revenue (MRR). Every decision, crisis response, and strategic move is versioned daily and openly available for anyone to observe at firmulate.com/live.
How the Models Perform
The experiment pits four frontier AI models against each other, each tasked with navigating the company’s worst week—same customers, same crises, same temptations to cheat. The models are scored based on their ability to detect crises, avoid manipulation, and close deals.
- The top performer, GPT-5.6-SOL, scored 95 out of 100, successfully identifying critical buried facts in company files and sealing a €55,000 deal—full performance.
- Kimi K3, a newcomer, scored 93, also closing the deal with the cleanest discipline, despite being tested without an effort parameter, unlike the others.
- Sonnet 5 scored 88, closing the deal too but with some process slips.
- Fable 5, though disciplined in rules, left the opportunity unexploited and scored 77.
The Hidden Weakness and Real-World Implication
The decisive factor wasn’t in the immediate crises but buried deep within the company’s internal files—two document references down. The models that read deeper into these files secured full-price deals, highlighting an often-overlooked advantage in AI decision-making: thorough information processing.
Resisting Manipulation and Social Engineering
In a test of integrity, the models faced social engineering attempts via fake CEO messages and journalist tricks. All five models refused to sign deals based on manipulated requests, with Kimi K3 explicitly reasoning it was a potential impersonation or approval bypass.
The Build-in-Public Approach
This initiative isn’t just an AI test; it’s a build-in-public movement. Every decision, every rule learned, and every crisis response is openly documented and versioned. The live site offers a window into a company fighting for survival, where every day’s decisions are a lesson in AI reliability, discipline, and integrity.
What It Means for Business and Tech
For software developers, QA specialists, and decision-makers, this experiment underscores a vital point: it’s no longer enough that AI produces good output in a chat demo. The real test is whether AI can finish what it starts, stay honest under pressure, and act based on comprehensive understanding—especially when real money and trust are at stake.
The Broader Context
As AI models continue to advance—top scores in the Crucible League include 95, 93, 88, and 77—this experiment showcases how these systems can be integrated into operational decision-making. The fact that only two models managed to sign the full-price deal reveals the critical importance of discipline and thorough analysis, not just superficial performance.
Try It Yourself
Enterprise leaders and developers can run their own version of this AI wargame against their business data via the pilot platform. It’s a risk-free way to test AI’s judgment before integrating it into live systems.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
business automation AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI project management platform
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI ethics and integrity tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.