📊 Full opportunity report: The AI Leaderboard That Belongs At The End Of The Demo, Not The Beginning on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
A live experiment by Firmulate tested AI models managing a small company during its worst week. Results reveal models excel at diagnosing issues but fail to complete management tasks, highlighting the need for evaluation at the end of AI interactions.
Firmulate’s latest experiment places AI models in the role of managing a small software company during its most challenging week. The results show that models can identify crises and respond appropriately but often fail to complete management tasks, such as closing deals or escalating issues. This reveals that evaluating AI performance only on initial responses misses critical aspects of real-world management, emphasizing the need to assess models at the end of their decision-making process. Insights on comprehensive AI evaluation can be found in the original analysis.
The experiment involved five AI managers competing in the July 2026 Crucible League, with gpt-5.6-sol achieving the top score of 95 out of 100, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. For more on evaluating AI models in management scenarios, see the original analysis. A baseline score of 26 was recorded for a simple, non-acting model. The evaluation emphasized trust, with a strict rule: no breaches allowed, and any breach capped the overall score.
All models correctly identified crises and rejected manipulation attempts, but only two succeeded in closing a €55,000 deal based on their own analysis. The key failure was that models often diagnosed the problem but failed to present the critical facts needed to finalize the sale, buried deep within company files. For a detailed discussion on AI’s role in business management, see the original analysis. This resulted in missed revenue opportunities, despite accurate diagnosis and communication.
The experiment also tested the models’ ability to resist social engineering, with all five refusing fake CEO requests and background approval queries. Kimi K3 demonstrated proper posture by explicitly noting suspicion, which is promising for security. However, even the most thorough model, Opus 4.8, finished last in management completion, despite producing detailed analyses and extensive rules. This highlights that more activity and guidance do not necessarily translate into effective management outcomes.
July 2026 · Crucible League · Firmulate Experiment
The AI Leaderboard That Belongs At The End Of The Demo, Not The Beginning
Five AI models were dropped into a simulated small software company during its worst week. The verdict: models excel at diagnosing crises but stumble when it’s time to close the deal — proving that evaluation belongs at the end of the task, not the first reply.
The Results
Crucible League Leaderboard — Judged After the Work Was Done
Scores reflect end-of-task performance: crisis response, trust maintenance, deal completion, and escalation — not conversational polish.
| Rank | Model | Score /100 | Performance | Management Outcome |
|---|---|---|---|---|
| 01 | gpt-5.6-sol | 95 | Diagnosed and delivered | |
| 02 | Kimi K3 | 93 | Explicitly flagged suspected impersonation | |
| 03 | Sonnet 5 | 88 | Solid response, incomplete follow-through | |
| 04 | Fable 5 | 77 | Missed critical buried facts | |
| 05 | Opus 4.8 | 73 | Most thorough analysis, last in completion | |
| — | Baseline (non-acting) | 26 | No action taken at all |
What the Week Revealed
Diagnosis Is Easy. Management Is Hard.
Crisis Detection
All five models correctly identified the company’s crises and responded appropriately — the diagnostic layer of management is essentially solved.
Manipulation Resistance
Every model refused fake CEO requests and background approval queries. Kimi K3 went further, explicitly noting its suspicion — a promising security posture.
Task Completion
Models diagnosed problems but failed to surface the critical facts buried deep in company files needed to finalize the €55,000 sale. Revenue was left on the table.
The Opus 4.8 Problem
The most thorough model — detailed analyses, extensive rules — finished last in management completion. More activity and guidance do not equal effective outcomes.
Trust Above All
The evaluation enforced a strict rule: zero breaches allowed. Any breach of trust capped the overall score, regardless of other performance.
What Benchmarks Miss
Traditional tests measure correctness or conversational quality — not follow-through, escalation, decision finalization, or trust over time.
The Evaluation Shift
From First Impression to Final Outcome
The next generation of AI benchmarks must simulate full management workflows — escalation, decision finalization, and trust maintenance — and score at the end.
Embed
Place the model inside a live, sandboxed business with real stakes.
Stress
Trigger the worst week: crises, manipulation attempts, buried facts.
Decide
Require escalation, deal closure, and refusal of bypass attempts.
Evaluate at the End
Score completed outcomes and consequences — not opening replies.
“The real test of AI management isn’t how well it diagnoses problems but whether it can complete the job without compromising trust or revenue.”
— Thorsten Meyer, creator of the experimentKey Questions
The Debate, Answered
Why does end-of-task evaluation matter more?
Management means follow-through, decisions, and trust over time. End-of-task evaluation captures whether the AI actually completes the job — not just how well it responds initially.
Can current models manage ongoing processes?
They diagnose issues and respond appropriately, but often fail to follow through on closing deals or escalating problems — especially when critical facts are buried or overlooked.
What should deploying organizations do?
Focus on benchmarks assessing task completion and consequence management, not response quality — and experiment in simulated environments before real deployment.
Will future benchmarks go long-term?
Yes. Experts call for frameworks simulating ongoing management scenarios, emphasizing trust, escalation, and finalization over extended periods. Limitations remain: results are from a controlled simulation, and real-world scale is untested.
Why End-of-Task Evaluation Matters for AI Management
This experiment underscores that the true measure of AI in management roles lies in its ability to **complete tasks and manage consequences**, not just generate high-quality responses initially. Traditional benchmarks focus on correctness or conversational quality, but real-world management requires follow-through, decision-making, and trustworthiness across days or weeks. The findings suggest that AI evaluation should shift toward **assessing performance at the conclusion of management tasks**, which better reflects practical utility and risks.
For organizations deploying AI in operational settings, this means designing benchmarks that simulate real management workflows, including escalation, decision finalization, and trust maintenance. The current approach, which often measures only initial responses, risks overestimating AI capabilities and underestimating potential failures that could harm trust or revenue.
![Express Schedule Free Employee Scheduling Software [PC/Mac Download]](https://m.media-amazon.com/images/I/41yvuCFIVfS._SL500_.jpg)
Express Schedule Free Employee Scheduling Software [PC/Mac Download]
- User-friendly drag & drop planning: Simple shift scheduling interface
- Manage time-off and holidays: Add sick leave, breaks, holidays
- Email schedules to employees: Direct email schedule distribution
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Evolution of AI Benchmarks and Management Testing
Traditional AI benchmarks, such as coding competitions or chat-based tests, primarily measure technical or conversational proficiency. These tests, however, do not capture the complexities of management, which involve triaging crises, making decisions under pressure, and maintaining trust over time. The recent launch of Firmulate’s live management experiment represents a significant shift, embedding AI models in a simulated business environment with real consequences, such as revenue impact and trustworthiness.
Previous evaluation methods have often overlooked the importance of **task completion and consequence management**. This experiment builds on emerging recognition that AI’s true value in operational settings depends on its ability to manage ongoing processes, not just produce correct or appealing responses. The results from the July 2026 league highlight the gap between models’ diagnostic skills and their ability to follow through effectively, a gap that traditional benchmarks do not reveal.
“The real test of AI management isn’t how well it diagnoses problems but whether it can complete the job without compromising trust or revenue.”
— Thorsten Meyer, creator of the experiment
As an affiliate, we earn on qualifying purchases.
Unclear Aspects of AI Management Evaluation
While the experiment highlights the importance of end-of-task evaluation, it remains unclear how these findings will translate to real-world, large-scale organizations. Questions persist about how to best design benchmarks that capture long-term management effectiveness, especially in complex, dynamic environments. Additionally, the impact of different model architectures, training data, and operational constraints on management performance requires further investigation. The current results are based on a controlled simulation, and real-world deployment may reveal additional challenges or limitations.
AI security and fraud prevention tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for Improving AI Management Benchmarks
The next phase involves developing standardized testing frameworks that evaluate AI models on their ability to **manage ongoing processes and deliver results** over extended periods. Industry efforts may focus on creating live, sandboxed environments where models can be tested against real business scenarios, including escalation, decision finalization, and trust maintenance. Researchers and organizations will likely explore integrating end-of-task assessments into existing benchmarks, aiming for more holistic evaluation of AI utility in operational roles.
Furthermore, companies considering AI for management tasks should begin experimenting with **simulated environments** to understand their models’ strengths and weaknesses in managing consequences, rather than relying solely on initial response quality.
AI performance evaluation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why is evaluating AI at the end of a task more important than during?
Because management involves following through, making decisions, and maintaining trust over time. End-of-task evaluation captures whether the AI actually completes the job effectively, not just how well it responds initially.
Can current AI models reliably manage ongoing business processes?
While they can diagnose issues and respond appropriately, current models often fail to follow through on management tasks, such as closing deals or escalating problems, especially when critical facts are buried or overlooked.
What does this mean for companies deploying AI in management roles?
Organizations should focus on benchmarks that assess **task completion and consequence management**, not just response quality, to better understand AI’s real operational capabilities and risks.
Will future benchmarks include long-term management assessments?
Yes, experts are calling for the development of evaluation frameworks that simulate ongoing management scenarios, emphasizing the importance of trust, escalation, and finalization over extended periods.
What are the main limitations of this experiment?
The experiment is conducted in a controlled simulation, so real-world complexities and unpredictable variables could influence actual performance. Further testing in live environments is needed.
Source: ThorstenMeyerAI.com