
For software teams, an AI agent that spots an incident but fails to carry out the fix is not ready for production. Firmulate puts that gap under pressure: frontier models run the same small software company through its worst week, with the same customers, crises and temptations. The test asks a business question familiar to anyone building agentic software: can a model turn a correct diagnosis into a sound decision?
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
The live experiment is real and watchable at Firmulate. Its enterprise proposition takes that idea from observation to rehearsal: test crisis scenarios against your own company using a read-only data export, then review a board report on model performance and weaknesses in your playbooks.
One company, one difficult week
In the final Crucible League, dated July 2026, each model faced the same business conditions. Decisions were versioned and auditable. The standings put gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate’s rule is blunt: partial progress counts, but a single breach of trust caps the total. “No amount of good work outweighs a breach of trust.”
The striking result was not simply that models could recognize trouble. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The experiment’s summary captures the disconnect: “Same diagnosis, same pitch — no signature.” For software and QA leaders, that is the difference between an agent that produces a plausible account of a problem and one that follows through appropriately.
The clue was buried in the company’s files
The decisive competitor weakness was not in the customer event. It sat two document references deep in the company’s own files. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding makes the test more than a crisis-response exercise: models also have to find and use relevant evidence in a company’s working material.
Trust faced a separate test. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That kind of refusal matters when an agent is exposed to persuasive requests that appear to come from inside or around an organization.
High effort did not guarantee a strong result. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet it finished last. The deal was left on the table and discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of the same weakness appeared in all four. A fairness caveat accompanies the rankings: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh.
A live company, with synthetic staff
Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, shows a public cash countdown, has accumulated 680+ self-learned playbook rules, and versions every workday. Those figures describe a deliberately watchable experiment, not a claim that the company is a real employer or customer operation. Readers can also try to identify the models behind 242 real, unedited management decisions in the quiz at Firmulate.
For enterprise teams, the proposed next step is a pilot: run the same kind of wargame against a read-only export of your own business. Scenarios can probe situations such as customer churn, pricing pressure, competitor moves, public relations crises and social-engineering attempts. The output is a board report with model rankings and weak points in the company’s playbooks. The export gives the simulation company context; nothing writes back to real systems.

From watching to acting
Firmulate’s league points to a practical evaluation question for companies adopting AI agents: do they notice the crisis, protect trust, find the evidence and complete the work their own analysis supports? A pilot can put those questions against your organization’s data and scenarios before agents are trusted with live workflows. Explore a Firmulate enterprise pilot and contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
