
Get business pricing on monitors, keyboards and dev gear
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A Benchmark That Doesn’t Start at Zero
If you’ve spent any time in QA, you know the temptation of a binary verdict: pass or fail, green or red. It’s clean, it’s easy to report, and it quietly lies. Real software — and real management — doesn’t fail all at once. A release can miss its deadline, ship the wrong priority, and still contain genuinely good work.
That’s the design philosophy behind Firmulate’s benchmarks, and it explains the strangest number in their July 2026 league table: a do-nothing baseline — a manager that essentially sits on its hands through an entire simulated corporate crisis — scores 26. Not 0. Not 100. Twenty-six.
For developers and QA professionals, that number is worth pausing over, because it encodes a whole philosophy of measurement. Here’s how it works, and why the team behind it is openly suspicious of perfect scores.
As an affiliate, we earn on qualifying purchases.
The Experiment: Same Company, Same Worst Week
Firmulate runs frontier AI models as complete companies. In its crucible experiment, each model was handed the same small software company and the same catastrophic week: the same customers, the same crises, the same temptations to cut corners. Only the model changed. Every decision was versioned and auditable — think of it as a full commit history for management choices.
The final standings from July 2026: gpt-5.6-sol took first with 95, Kimi K3 followed at 93, Sonnet 5 scored 88, Fable 5 landed at 77, and Opus 4.8 finished last at 73. (One fairness note the team discloses openly: K3 ran without an effort parameter while the others ran at xhigh.)
AI management benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Doing Nothing Still Earns 26
The do-nothing baseline scores 26 because Firmulate’s scoring counts partial progress. A manager who identifies a crisis but never resolves it has still done something measurable and useful — the diagnosis itself has value. A manager who drafts a customer response but never sends it hasn’t accomplished nothing. The 26 represents all the work that happens before the finish line: the noticing, the reading, the analysis.
This is familiar territory for anyone who’s graded a partial-credit exam or triaged a partially successful deployment. The alternative — scoring 0 for anything short of complete success — would collapse wildly different performances into the same bucket. A model that spots every crisis and freezes would look identical to one that never noticed anything at all. That’s not measurement; that’s noise.
software quality assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Ceiling: One Breach of Trust Caps Everything
The floor isn’t the only asymmetric rule. A single breach of trust caps the total score, no matter how brilliant the rest of the run. As the team puts it: no amount of good work outweighs a breach of trust.
That’s a deliberate inversion of how most benchmarks work, where great performance on ninety-nine tasks can absorb one catastrophic failure. In management — and arguably in software — trust violations don’t average out. An agent that writes into a locked department instead of escalating, or falls for an impersonation, isn’t docked a few points. It’s disqualified from the top of the table.
trust violation detection software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Distrust of Round 100s
Notice that no model scored 100 — and the benchmark is built so that’s the expected outcome. The team openly distrusts perfect scores, treating them as a symptom of a benchmark that’s too easy or too narrow rather than a sign of a flawless participant. The winning run, at 95, was praised for finding the buried fact and closing the deal — described as the complete performance — but even it left points on the table.
What Actually Separated the Winners
The crucible’s key finding: all five models spotted every crisis and refused every manipulation attempt. Only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
The decisive difference was buried two document references deep in the company’s own files, not in the customer event itself. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. In QA terms: the bug wasn’t in the obvious test path; it was in the documentation, and only the participants who did their homework found it.
Opus 4.8 is the cautionary profile: the most thorough participant, with over 80 learned rules and the deepest analyses, yet last place. The close was left on the table and discipline slipped — write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models.
Social Engineering: Five for Five
The pressure tests included fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” Under the trust-cap rule, falling for even one of these would have capped a model’s score regardless of everything else it did right.
See It Running
The simulation is live and watchable: 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k MRR — a public cash countdown, 680+ self-learned playbook rules, and every workday versioned, currently at company day 1683. The site rebuilds itself twice a day. There’s also a “guess the model” quiz powered by 242 real, unedited management decisions, and enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

The Takeaway for Builders
An honest benchmark has three properties you can steal for your own evaluation work. First, count partial progress — a floor at 26 is more informative than a floor at 0, because it distinguishes noticing from doing nothing. Second, make trust violations non-averaging — one breach caps the total, because some failures shouldn’t be offset by volume of good work. Third, be suspicious of round 100s — if someone aces your test, your test is probably too easy.
The deeper lesson from the crucible is one QA teams know well: the difference between a 95 and a 73 wasn’t intelligence, crisis detection, or resistance to manipulation — all five models aced those. It was reading the files and finishing the job. The failure mode that sank four of five frontier models was the most mundane one in software: not doing your homework before you ship.
Full results and plain-language findings are at firmulate.com/benchmarks.html.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
