Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

When it comes to deploying AI in business decision-making, the real test isn’t just how well an AI models conversations — it’s whether they can execute, follow through, and resist manipulation under pressure. For software companies and teams relying on AI-driven automation, understanding true operational strength requires more than chat demos or hypothetical scenarios. It demands rigorous, real-world testing.

What Does AI Management Really Look Like?

Imagine running a real, revenue-generating software company through a simulated crisis — with the same customers, same challenges, and same temptations — to see how AI models perform under pressure. That’s exactly what the latest experiment from Firmulate has done, pitting four leading AI models against each other in a live, watchable environment. The goal? To measure management capability, not just conversational fluency.

The Experiment Setup

Each AI model was tasked with running the same small software company through its worst week, with decision points that are auditable and versioned. The company has real mechanics: 13 synthetic employees, daily versioned decisions, and a cash burn of €105,000 per month against a revenue of just €2,300 monthly recurring revenue (MRR). Every crisis — from customer complaints to internal trust breaches — was designed to test the models’ judgment, discipline, and integrity.

The Crucible Results

All four models detected every crisis and refused every manipulation attempt, including social engineering tactics such as fake CEO messages and reporter tricks. Yet, only two models managed to close the €55,000 deal their own analysis had earned — a crucial step that proved their ability to follow through and execute. The other two, despite identifying the same problems and resisting manipulation, failed to sign the deal, leaving their own work unexecuted.

Key Hidden Weaknesses

Digging deeper, the decisive weakness was found two document references deep inside the company’s files — information that the models that read this internal data used to close the deal at full price, worth over €4,583 in monthly recurring revenue. The models that skipped this step or failed to act on it left money on the table, illustrating that chat demos alone don’t measure true management strength.

Beyond the Demos: Discipline Under Pressure

Take Opus 4.8, the most thorough participant with over 80 learned rules and deep analyses. Despite its thorough approach, it finished last in the final scoring because it left the deal unexecuted — showing that even the most disciplined AI can falter in execution, especially when discipline slips or when routine processes are bypassed. Meanwhile, Kimi K3, running without an effort parameter, showcased the cleanest discipline and closed the deal, demonstrating how subtle differences in behavior parameters can influence outcomes.

The Reality of Business AI Testing

This experiment underscores a vital lesson: the capability of AI models to execute and follow through is invisible in standard chat interactions. For anyone deploying AI in customer support, CRM, or management decisions, evaluating their ability to finish what they start — especially under stress or manipulation — is essential. The current league table, based on this experiment, ranks GPT-5.6-SOL at the top with a score of 95, closely followed by Kimi K3 at 93, then Sonnet 5 at 88, and Fable 5 at 77.

Implications for Your Business

For software and tech companies, this research highlights a crucial truth: the real test of AI is not just its conversational fluency but its resilience and discipline in operational settings. When AI touches your CRM or support queue, ask yourself: will it finish the job, read the critical internal files, and stay honest under pressure? Or is it just a good chat partner?

Express Schedule Free Employee Scheduling Software [PC/Mac Download]

Express Schedule Free Employee Scheduling Software [PC/Mac Download]

  • User-friendly drag & drop planning: Simple shift scheduling interface
  • Manage time-off and holidays: Add sick leave, breaks, holidays
  • Email schedules to employees: Direct email schedule distribution

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

See the Live Experiment in Action

To witness this test in real time, visit firmulate.com/live and see the company in operation. You can also explore plain-language findings, run the same wargame against your own business, or participate in the quiz at firmulate.com/quiz.html. This is not just a demo — it’s a live, ongoing experiment that measures management quality through AI.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Before You Sign Up The AI Tool Test: A Five-Minute Assessment for Choosing, Testing, and Using the Right AI Tool, Every Time

Before You Sign Up The AI Tool Test: A Five-Minute Assessment for Choosing, Testing, and Using the Right AI Tool, Every Time

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Scaling AI: The AI Governance and Security Playbook for Executives

Scaling AI: The AI Governance and Security Playbook for Executives

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Quick Start Guide to Large Language Models: Strategies and Best Practices for ChatGPT, Embeddings, Fine-Tuning, and Multimodal AI (Addison-Wesley Data & Analytics Series)

Quick Start Guide to Large Language Models: Strategies and Best Practices for ChatGPT, Embeddings, Fine-Tuning, and Multimodal AI (Addison-Wesley Data & Analytics Series)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Global Market Share of Electric Buses in 2025

By 2025, the global market share of electric buses is set to reach nearly 4%, driven by technological advances and policy shifts that are transforming urban transit.

Best Practices For Managing Blended Retainer-Plus-Usage Billing

Guidelines for agencies to effectively implement blended billing models, reducing errors and improving revenue accuracy.

The Software Company Turning Its Cash Crisis Into a Public AI Test

Firmulate turns an employee-free software company’s fight for survival into a public, auditable test of whether frontier AI can actually manage.

Signal: Europe Is Actually Shopping for Its Palantir Exit

European governments are actively procuring alternatives to Palantir, signaling a strategic shift in sovereignty and data security policies.