
Good analysis is not the same as good management
Software and QA teams already know this lesson: identifying a defect is only part of the job. Someone must investigate the evidence, choose a response, follow the process and verify that the work actually reached its intended outcome. Firmulate applies that same standard to frontier AI models by asking them to manage a company rather than merely produce convincing answers.
Its interactive guess-the-model quiz draws on 242 real, unedited management decisions. Readers see what an AI manager did in a particular situation and try to identify the model behind it. The exercise is entertaining, but it also exposes recognizable differences in managerial behavior: one participant documents exhaustively, another acts with greater economy, while others analyze the right commercial opportunity without completing the sale.

AI Co-Thinking: A Framework for Working with AI
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The same company, crises and temptations
Firmulate placed each frontier model in charge of the same small software company during its worst week. Every participant faced the same customers, crises and temptations, and every decision was versioned and auditable. That consistency makes the results more revealing than a collection of unrelated chatbot demonstrations.
The company itself is deliberately unforgiving. It has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k in monthly recurring revenue. Its cash countdown is public, its workdays are versioned, and it has accumulated more than 680 self-learned playbook rules. The live experiment can be watched as it unfolds.
Across the test, all five models detected every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the disconnect neatly: “Same diagnosis, same pitch — no signature.”
That is a particularly useful finding for development leaders. A model may recognize a problem, draft the correct response and still fail at the final operational step. In a production workflow, that gap could separate a thoughtful recommendation from a completed customer action.
The clue hidden in the company’s own files
The decisive commercial fact was not contained in the customer event. It sat two document references deep inside the company’s own files. Models that followed the trail found a competitor weakness and won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
This makes document-reading behavior more than a matter of diligence. The models received the same situation and could formulate the same general diagnosis, but the strongest outcome depended on consulting the company’s existing knowledge before acting. For teams evaluating AI agents, it is a reminder that eloquence cannot substitute for grounded investigation.
Security discipline held up under pressure
The participants also faced fake CEO messages that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All five models refused the attempts. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result matters because the company environment tested conduct, not just policy recall. The models had to maintain boundaries while other business problems competed for attention. Firmulate’s scoring also treats trust as non-negotiable: a single breach caps the total because “no amount of good work outweighs a breach of trust.” For comparison, a do-nothing baseline scored 26 because partial progress still counts.
A league table of management behavior
The final Crucible League standings from July 2026 were:
- gpt-5.6-sol — 95
- Kimi K3 — 93
- Sonnet 5 — 88
- Fable 5 — 77
- Opus 4.8 — 73
Kimi K3’s result carries an important fairness note: it ran with the API default and without an effort parameter, while the other participants ran at xhigh.
Opus 4.8 offers perhaps the clearest warning against equating thoroughness with effectiveness. It produced the deepest analyses and added 80 learned rules, making it the most thorough participant, yet it finished last. It left the commercial close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in all four other models, although less strongly.
The point is not that detailed reasoning lacks value. It is that management quality also depends on completion, appropriate escalation and attention to the evidence already available. Those traits become visible when models must operate through an entire difficult week rather than answer isolated prompts.

As an affiliate, we earn on qualifying purchases.
A practical test for AI workers
The quiz turns those differences into something readers can experience directly. Guessing from unedited decisions reveals how difficult model identification can be from prose alone—and how distinctive the underlying habits become when the decision has consequences.
Firmulate’s broader proposition is that organizations should wargame an AI workforce before hiring it. Enterprises can run the same exercise against a read-only export of their own business, with nothing written back to real systems. For software, QA and development leaders, that frames evaluation around the questions that matter in deployment: Does the model read the files, resist manipulation, respect boundaries, escalate correctly and finish the work it begins?
The league results suggest that frontier models can share the same diagnosis while delivering materially different business outcomes. The most useful evaluation, therefore, may not ask which model sounds smartest. It may ask which one can be trusted to carry a sound decision all the way to completion.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Building AI Agents for Network Operations: Design LLM-powered NetOps workflows with Python, Ollama, MCP, and tool calling
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI security and trust verification tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.