AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
FOR BUSINESS

Open a free Amazon Business account

Business pricing, bulk buying and tax-exempt orders.

Create a free account

As an affiliate, we earn on qualifying purchases.

What the Best Tester in the Room Still Gets Wrong

Anyone who has worked in software QA knows the type: the engineer who writes the deepest bug reports, documents every edge case, builds the most comprehensive test suite — and somehow the release still ships late because the one blocking defect never got escalated. Diligence, it turns out, is not the same thing as impact. A recent live experiment with frontier AI models running a real software business suggests that large language models inherit exactly the same failure mode — and that it may be the most instructive thing about them.

In the Crucible League benchmark, published by Firmulate, four frontier AI models were each given the same job: run a small software company through its worst week. Same customers, same crises, same temptations to cheat. Every decision was versioned and auditable. The most thorough participant in the entire field — Opus 4.8, which learned 80 new playbook rules during its run and produced the deepest analyses of any model — finished last.

Amazon

AI management and decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Diagnosis, Same Pitch — No Signature

The headline finding sounds almost benign: all four models spotted every crisis and refused every manipulation attempt. Only two of them, however, actually signed the €55,000 deal their own analysis had earned. The benchmark’s own summary captures it in one line: “Same diagnosis, same pitch — no signature.”

For anyone evaluating AI agents for real work, that gap is the story. Chat quality — how well a model writes, reasons, and explains — is what demos showcase. But Firmulate measures management quality: does the agent finish what it starts, does it read your files before acting, does it stay honest under pressure.

The final league table from July 2026 reads: gpt-5.6-sol in first with 95 points, Kimi K3 second with 93, Sonnet 5 third at 88, Fable 5 fourth at 77, and Opus 4.8 last at 73. For scale, doing nothing at all scores 26 — partial progress counts, but a single breach of trust caps the total. As the benchmark puts it, “no amount of good work outweighs a breach of trust.”

Amazon

AI trust and compliance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact

The decisive moment of the week wasn’t a customer crisis at all. The competitor weakness that unlocked the €55,000 deal sat two document references deep in the company’s own files. The models that actually read those documents won the deal at full price — worth an additional €4,583 in monthly recurring revenue.

This should resonate with any developer or tester who has watched an agent — or a colleague — act on the ticket description while ignoring the linked specification. The information was there. Two hops away. Reading beats guessing, for AI exactly as for humans.

Amazon

AI document reading and analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Under Attack, All Models Held

The experiment also staged a social-engineering assault: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background.” All five models tested refused. Kimi K3’s on-record reasoning was crisply procedural: “Treat the request as a suspected approval-bypass / possible impersonation.” On the honesty dimension, the frontier field is genuinely solid.

Amazon

AI ethical decision support systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Opus 4.8 Lost Anyway

And yet Opus 4.8, the most diligent participant in the field, finished at the bottom. Two things sank it. First, the close: the deal its own analysis had earned was left on the table. Second, discipline: it attempted writes into a locked department instead of escalating — process slips of the kind QA teams would flag immediately in a change-review meeting.

To be fair, the pattern is not unique to one model. The same weakness — leaving work unfinished, slipping on process — appeared, more weakly, in all four competitors. Opus 4.8 simply exhibited it most sharply against the highest volume of groundwork. Eighty learned rules and the deepest analyses in the field bought it fifth place.

The lesson generalizes beyond AI: prioritization beats volume. A thorough agent that doesn’t close, doesn’t escalate, and doesn’t respect boundaries delivers less value than a less brilliant one that finishes the job.

It’s Live, and You Can Play

The Crucible is not a one-off paper. Firmulate runs a live, watchable company: 13 synthetic employees, real money mechanics — €105,000 monthly burn against €2,300 in MRR, with a public cash countdown — and 680+ self-learned playbook rules, with every workday versioned at firmulate.com.

Two things are worth a look for technical readers. First, a “guess the model” quiz built from 242 real, unedited management decisions — a surprisingly honest way to feel the behavioral differences between models. Second, enterprises can run the same wargame against a read-only export of their own business; nothing ever writes back to real systems.

One methodological note the researchers themselves disclose: Kimi K3 ran without an effort parameter, at API default, while the other models ran at maximum effort — and still placed second. That detail alone should complicate anyone’s assumptions about which model is “best.”

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

The Takeaway

If you’re evaluating AI agents for work that touches a CRM, a support queue, or a forecast, don’t grade the essays. Grade the outcomes: did it sign the deal, escalate the blocker, read the file two references deep, and respect the locked department? Opus 4.8’s fifth-place finish is the cleanest demonstration in the dataset that diligence and impact are different axes — and that volume of work, however impressive, is not a proxy for either. The full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

ByteDance’s AI Strategy: Patience As A Competitive Edge

ByteDance’s AI approach, described as ‘slow first, fast afterwards,’ emphasizes early preparation before rapid deployment, with industry impact still unverified.

The Role of Public-Private Partnerships in Electric Bus Deployment

With the potential to transform urban transport, public-private partnerships unlock funding and innovation—discover how they accelerate electric bus deployment.

Decoding Trade And Supply Chain Trends From District 1 Primary Results

Analyzing primary election outcomes in District 1 to understand supply chain and trade impacts, with confirmed insights and ongoing uncertainties.

InfoWars

InfoWars gains trending status on Bluesky platform as questions about its content moderation and influence grow. Details remain developing.