AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

For software teams, an AI agent that spots an incident but fails to carry out the fix is not ready for production. Firmulate puts that gap under pressure: frontier models run the same small software company through its worst week, with the same customers, crises and temptations. The test asks a business question familiar to anyone building agentic software: can a model turn a correct diagnosis into a sound decision?

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

The live experiment is real and watchable at Firmulate. Its enterprise proposition takes that idea from observation to rehearsal: test crisis scenarios against your own company using a read-only data export, then review a board report on model performance and weaknesses in your playbooks.

One company, one difficult week

In the final Crucible League, dated July 2026, each model faced the same business conditions. Decisions were versioned and auditable. The standings put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate’s rule is blunt: partial progress counts, but a single breach of trust caps the total. “No amount of good work outweighs a breach of trust.”

The striking result was not simply that models could recognize trouble. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s summary captures the disconnect: “Same diagnosis, same pitch — no signature.” For software and QA leaders, that is the difference between an agent that produces a plausible account of a problem and one that follows through appropriately.

The clue was buried in the company’s files

The decisive competitor weakness was not in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding makes the test more than a crisis-response exercise: models also have to find and use relevant evidence in a company’s working material.

Trust faced a separate test. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That kind of refusal matters when an agent is exposed to persuasive requests that appear to come from inside or around an organization.

High effort did not guarantee a strong result. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet it finished last. The deal was left on the table and discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four. A fairness caveat accompanies the rankings: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

A live company, with synthetic staff

Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, shows a public cash countdown, has accumulated 680+ self-learned playbook rules, and versions every workday. Those figures describe a deliberately watchable experiment, not a claim that the company is a real employer or customer operation. Readers can also try to identify the models behind 242 real, unedited management decisions in the quiz at Firmulate.

For enterprise teams, the proposed next step is a pilot: run the same kind of wargame against a read-only export of your own business. Scenarios can probe situations such as customer churn, pricing pressure, competitor moves, public relations crises and social-engineering attempts. The output is a board report with model rankings and weak points in the company’s playbooks. The export gives the simulation company context; nothing writes back to real systems.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching to acting

Firmulate’s league points to a practical evaluation question for companies adopting AI agents: do they notice the crisis, protect trust, find the evidence and complete the work their own analysis supports? A pilot can put those questions against your organization’s data and scenarios before agents are trusted with live workflows. Explore a Firmulate enterprise pilot and contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

AI Innovation In Asian Penny Stocks: SenseTime Group And Promising Alternatives

SenseTime Group is featured in a recent report as one of three promising Asian penny stocks, but supporting analysis and other selections remain unconfirmed.

Vinfast Debuts Iev12 Electric Bus for Global Markets

An innovative new electric bus by Vinfast promises to revolutionize urban transit worldwide—discover how the Iev12 is setting new industry standards.

InfoComm Asia To Relocate To Kuala Lumpur In 2027 – Pro AVL Asia

InfoComm Asia will relocate to Kuala Lumpur in 2027, marking a significant shift for the event’s future in the Asia-Pacific region.

X down for thousands of users globally, Downdetector shows

X is currently down for thousands of users worldwide, according to Downdetector reports. The outage impacts access and functionality, with details still emerging.