AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
Live on firmulate.com.

When it comes to deploying AI in business decision-making, the real test isn’t just how well an AI models conversations — it’s whether they can execute, follow through, and resist manipulation under pressure. For software companies and teams relying on AI-driven automation, understanding true operational strength requires more than chat demos or hypothetical scenarios. It demands rigorous, real-world testing.

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

What Does AI Management Really Look Like?

Imagine running a real, revenue-generating software company through a simulated crisis — with the same customers, same challenges, and same temptations — to see how AI models perform under pressure. That’s exactly what the latest experiment from Firmulate has done, pitting four leading AI models against each other in a live, watchable environment. The goal? To measure management capability, not just conversational fluency.

The Experiment Setup

Each AI model was tasked with running the same small software company through its worst week, with decision points that are auditable and versioned. The company has real mechanics: 13 synthetic employees, daily versioned decisions, and a cash burn of €105,000 per month against a revenue of just €2,300 monthly recurring revenue (MRR). Every crisis — from customer complaints to internal trust breaches — was designed to test the models’ judgment, discipline, and integrity.

The Crucible Results

All four models detected every crisis and refused every manipulation attempt, including social engineering tactics such as fake CEO messages and reporter tricks. Yet, only two models managed to close the €55,000 deal their own analysis had earned — a crucial step that proved their ability to follow through and execute. The other two, despite identifying the same problems and resisting manipulation, failed to sign the deal, leaving their own work unexecuted.

Key Hidden Weaknesses

Digging deeper, the decisive weakness was found two document references deep inside the company’s files — information that the models that read this internal data used to close the deal at full price, worth over €4,583 in monthly recurring revenue. The models that skipped this step or failed to act on it left money on the table, illustrating that chat demos alone don’t measure true management strength.

Beyond the Demos: Discipline Under Pressure

Take Opus 4.8, the most thorough participant with over 80 learned rules and deep analyses. Despite its thorough approach, it finished last in the final scoring because it left the deal unexecuted — showing that even the most disciplined AI can falter in execution, especially when discipline slips or when routine processes are bypassed. Meanwhile, Kimi K3, running without an effort parameter, showcased the cleanest discipline and closed the deal, demonstrating how subtle differences in behavior parameters can influence outcomes.

The Reality of Business AI Testing

This experiment underscores a vital lesson: the capability of AI models to execute and follow through is invisible in standard chat interactions. For anyone deploying AI in customer support, CRM, or management decisions, evaluating their ability to finish what they start — especially under stress or manipulation — is essential. The current league table, based on this experiment, ranks GPT-5.6-SOL at the top with a score of 95, closely followed by Kimi K3 at 93, then Sonnet 5 at 88, and Fable 5 at 77.

Implications for Your Business

For software and tech companies, this research highlights a crucial truth: the real test of AI is not just its conversational fluency but its resilience and discipline in operational settings. When AI touches your CRM or support queue, ask yourself: will it finish the job, read the critical internal files, and stay honest under pressure? Or is it just a good chat partner?

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

See the Live Experiment in Action

To witness this test in real time, visit firmulate.com/live and see the company in operation. You can also explore plain-language findings, run the same wargame against your own business, or participate in the quiz at firmulate.com/quiz.html. This is not just a demo — it’s a live, ongoing experiment that measures management quality through AI.

Infographic — Four AI Models Ran the Same Company Through Its Worst Week. Only Two Finished the Job.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision-making testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI operational discipline assessment

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI performance evaluation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Labor Displacement Data: What Q1-Q2 2026 Actually Shows

This report examines confirmed labor market shifts in early 2026, highlighting structural impacts of AI-driven automation on specific worker groups.

Isar Aerospace Reaches Orbit And Deploys Payloads On Second Flight

German launch firm Isar Aerospace successfully reached orbit and deployed payloads on its second mission, marking a significant milestone in its development.

Behind Xbox’s Big Layoffs, a Streaming Strategy That Failed

Microsoft’s recent layoffs at Xbox are linked to the company’s unsuccessful streaming service strategy, confirmed by sources and industry analysis.

Fair-value appraisals for used GPUs and AI hardware

A new manual valuation approach for used GPUs and AI hardware aims to create transparent, reliable pricing benchmarks for brokers in the secondary market.