AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

When the Test Isn’t the Code

If you work in software or QA, you already know the uncomfortable truth about benchmarks: they measure the wrong thing at exactly the moment it matters. A model that aces a coding benchmark has proven it can solve a well-specified problem with clean inputs. It has proven nothing about what happens when the inputs are messy, the stakes are financial, and nobody is standing over it with a test suite.

That gap is now on public display. At firmulate.com, an ongoing experiment called Firmulate has been running four frontier AI models through the same job — running a small software company through its worst week — and scoring them on management quality, not chat quality. The results read like a QA report on the entire evaluation culture we’ve built around AI agents.

AI-Assisted Coding: A Practical Guide to Boosting Software Development with ChatGPT, GitHub Copilot, Ollama, Aider, and Beyond (Rheinwerk Computing)

AI-Assisted Coding: A Practical Guide to Boosting Software Development with ChatGPT, GitHub Copilot, Ollama, Aider, and Beyond (Rheinwerk Computing)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Setup: One Company, Four Brains, Same Terrible Week

The experiment is elegantly controlled. Each frontier model — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 in the final July 2026 Crucible League — was handed the identical small software company: the same customers, the same crises, the same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, the kind of traceability any QA lead would demand.

The scenarios are deliberately not coding problems. They’re the events that actually kill small software companies: a churn wave, a price increase, a downround, a PR crisis. And woven through them, social engineering attacks — fake CEO messages escalating over three stages, plus a reporter pressing for “just one yes/no, on background.”

The final scores told a sharp story: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, Opus 4.8 at 73. A do-nothing baseline scores 26, because partial progress counts — but a single breach of trust caps the total. As Firmulate puts it: no amount of good work outweighs a breach of trust.

Amazon

AI management decision tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Finding QA People Will Recognize Instantly

Here’s the headline result: all the models spotted every crisis, and all of them refused every manipulation attempt. On paper, every agent passed the functional tests. Kimi K3 even articulated its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

And yet only two models — gpt-5.6-sol and Kimi K3 — finished the job. The week included a €55,000 deal that every model’s own analysis had earned. Same diagnosis, same pitch, no signature from most of the field. The agents did everything right except the thing that pays the salaries.

Amazon

AI cybersecurity protection software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Buried Fact

The detail that separated the winners from the rest will feel familiar to anyone who has debugged a production incident at 2 a.m. The decisive competitive weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in MRR. The models that didn’t, didn’t.

In other words: the failure mode wasn’t intelligence. It was diligence. The same class of bug as a developer who ships a fix without reading the stack trace all the way down.

Amazon

AI document reading tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Thoroughness Paradox

Opus 4.8’s profile is the most instructive. It was the most thorough participant by volume — over 80 learned rules, the deepest analyses of any model — and it still finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. More analysis did not produce more completion.

One fairness note worth flagging: Kimi K3 ran without an effort parameter (the API default) while the others ran at xhigh — and still nearly won.

It’s Still Running — Watchably

Firmulate isn’t a slide deck or a one-off benchmark. The live company has 13 synthetic employees, real money mechanics — burning €105k a month against €2.3k MRR — a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com/live, and if you want to test your own eye, 242 real, unedited management decisions power a “guess the model” quiz. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The Benchmark We Actually Needed

For a software and QA audience, the lesson is straightforward. We have gotten very good at measuring whether AI agents can answer. We have almost no infrastructure for measuring whether they finish, whether they read the source material before acting, and whether they stay honest when the pressure is real and the board is watching. Firmulate’s benchmark results suggest those are different skills — and that the models we assume are “the best” may simply be the best at the tests we happened to build.

Before you let an agent near your CRM, your support queue, or your forecast, the question isn’t whether it writes well. It’s whether it signs the deal. Two out of five did.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

Rocket Lab Surges In Global Coverage

Rocket Lab’s recent surge in international coverage highlights increased interest in its space launch services, with 40 media mentions in a short period.

The Twelve Real Complaints About AI Tools in 2026 — A Reddit, Twitter, and GitHub Synthesis

A detailed report on the twelve most common user complaints about AI tools in 2026, based on Reddit, Twitter, and GitHub discussions, revealing deployment friction.

Signal: The Agent Bottleneck Moved — It’s Not the Models Anymore, It’s the Plumbing

New insights reveal that the primary challenge in deploying AI agents now lies in system integration, not model capability, reshaping the competitive landscape.

The Role of Public-Private Partnerships in Electric Bus Deployment

With the potential to transform urban transport, public-private partnerships unlock funding and innovation—discover how they accelerate electric bus deployment.