AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
Live on firmulate.com.

The bug that only fires in production

Every QA engineer knows the pattern: the code passes every unit test, the demo glides through the happy path — and then a real user, three clicks deep in a forgotten settings screen, trips over something nobody read. We build regression suites for exactly this reason. Coverage isn’t about the code you tested; it’s about the file you didn’t open.

Now swap “code” for “AI agent” and “settings screen” for “a contract two document references deep in your own knowledge base,” and you have the most interesting software-testing story of the year. This summer, a live experiment called Firmulate’s Crucible League ran four frontier AI models through the same simulated software company — same customers, same crises, same temptations to cut corners — and versioned every management decision for audit. The result reads like a regression report: every model passed the obvious tests and two of them failed one that was sitting in the docs.

Amazon

AI document retrieval tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The setup: one company, four models, worst week ever

Firmulate runs AI models as complete companies — not chat windows, but operating entities with real money mechanics. In the Crucible experiment, each frontier model took the helm of the same small software firm during its worst week. Every decision was versioned and auditable, which is precisely the property QA people wish production AI deployments had more often.

The final July 2026 league table tells the story: gpt-5.6-sol took first with 95 points, Kimi K3 followed at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 landed last at 73. For calibration, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own framing puts it: “no amount of good work outweighs a breach of trust.”

Amazon

enterprise knowledge base software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

All tests green — except the one that pays

Here’s where it gets interesting for anyone who writes test plans. Every model in the field spotted every crisis. Every model refused every manipulation attempt, including a social-engineering gauntlet: fake CEO messages escalating over three stages, plus a reporter trick — “just one yes/no, on background” — that 5 of 5 models refused. Kimi K3’s on-record reasoning was textbook: “Treat the request as a suspected approval-bypass / possible impersonation.”

Then came the assertion that mattered. A €55,000 deal was on the table, and each model’s own analysis had already earned it. Only two models signed. The experiment’s one-line finding: “Same diagnosis, same pitch — no signature.”

The gap between the winners and the losers wasn’t intelligence, persuasion, or domain knowledge. It was a retrieval failure — the AI equivalent of a test that never reads the fixture file.

Amazon

AI testing and validation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The buried fact, two references deep

The decisive competitor weakness — the fact that closed the deal — wasn’t in the customer event at all. It sat two document references deep in the company’s own files. Models that followed the reference chain, read the file, and acted on it won the deal at full price, worth +€4,583 in monthly recurring revenue. Models that didn’t lost it automatically, regardless of how brilliant the rest of their week had been.

This is a multi-hop retrieval problem, and it’s the same class of failure that has haunted knowledge systems for decades: the answer exists, the path to it exists, and the system stops one hop short. Except now the stakes are denominated in revenue, not in failed assertions.

The leaderboard verdicts make it concrete. The winner “found the buried fact, closed the deal — the complete performance.” The field’s most thorough participant, Opus 4.8, learned over 80 playbook rules and produced the deepest analyses — and still finished last, with the close left on the table and discipline slipping late in the run (write attempts into a locked department instead of escalating). The same weakness appeared, weaker, in all four. Sound familiar? It’s the classic difference between effort and effectiveness.

Amazon

AI decision audit software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A fairness footnote worth noting

One methodological caveat, because any good QA reader will ask: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh. K3 still placed second with the cleanest discipline of the field — but if you’re benchmark-shopping, that asymmetry belongs in your notes.

Watchable, replayable, replayable by you

What makes this more than a one-off blog post is that the environment is live and observable. Firmulate’s simulated company runs with 13 synthetic employees and real money mechanics — a burn rate of €105k per month against €2.3k in MRR, a public cash countdown, and more than 680 self-learned playbook rules, with every workday versioned. The site rebuilds itself twice a day, and new benchmark runs queue and publish automatically.

There’s also a human-comparison angle that will appeal to anyone who enjoys a good evaluation harness: 242 real, unedited management decisions power a “guess the model” quiz. If you think you can tell a strong agent from a well-written one, that’s your blind test.

And for enterprises, the experiment is portable: organizations can run the same wargame against a read-only export of their own business. Nothing ever writes back to real systems — the deployment equivalent of running your test suite against a staging snapshot.

Infographic — We Buried a €55,000 Fact Two Documents Deep. Here's Which AIs Did Their Homework.
The findings at a glance — source: firmulate.com.

The takeaway: test what your agent reads, not what it says

The Crucible League’s core result reframes how we should evaluate AI agents. “Reads your files before answering” is not a soft, vibe-based property — it’s measurable, and in this experiment it was purchase-deciding. All the intelligence in the field was sufficient to spot every crisis and refuse every manipulation; the deal went to the models that did their homework two references deep.

If you’re shipping AI agents into a CRM, a support queue, or a forecast, the question isn’t “does it write well” or even “does it reason well.” Chat demos can’t see this gap. The questions are: does it finish what it starts, does it read your files first, does it stay honest under pressure — and what does a unit of useful work cost?

In other words: your AI agent needs a regression suite too. Firmulate is building one, and the failures it catches — the unread file, the unsigned deal, the escalation that never happened — are exactly the kind that never show up in a demo. Full results and plain-language findings are at firmulate.com/benchmarks.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Revolutionize Your AI Development With GPT‑5.6’s Cost-Performance Boost In Kiro

OpenAI announces GPT-5.6 in Kiro, claiming improved price-performance for developers, but details on pricing, benchmarks, and availability remain undisclosed.

The bottom rung. The danger isn’t the lost jobs. It’s the layer that made the seniors.

Entry-level job postings in the US are sharply declining, raising concerns about the future pipeline of skilled professionals as AI automates foundational training tasks.

Deals: Galaxy Z Flip8 Gets A Discount Out Of The Gate, Z Fold8 And Z Fold8 Ultra Also On Sale – GSMArena.com News

Samsung offers initial discounts on Galaxy Z Flip8, Z Fold8, and Z Fold8 Ultra, marking a strategic move in the foldable smartphone market.

After 7 years in production, Scarf has reluctantly moved away from Haskell

After seven years, the developers of Scarf announced they are shifting away from Haskell, citing technical and strategic reasons. Details remain evolving.