AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: The AI Leaderboard That Belongs At The End Of The Demo, Not The Beginning on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

A live experiment by Firmulate tested AI models managing a small company during its worst week. Results reveal models excel at diagnosing issues but fail to complete management tasks, highlighting the need for evaluation at the end of AI interactions.

Firmulate’s latest experiment places AI models in the role of managing a small software company during its most challenging week. The results show that models can identify crises and respond appropriately but often fail to complete management tasks, such as closing deals or escalating issues. This reveals that evaluating AI performance only on initial responses misses critical aspects of real-world management, emphasizing the need to assess models at the end of their decision-making process. Insights on comprehensive AI evaluation can be found in the original analysis.

The experiment involved five AI managers competing in the July 2026 Crucible League, with gpt-5.6-sol achieving the top score of 95 out of 100, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. For more on evaluating AI models in management scenarios, see the original analysis. A baseline score of 26 was recorded for a simple, non-acting model. The evaluation emphasized trust, with a strict rule: no breaches allowed, and any breach capped the overall score.

All models correctly identified crises and rejected manipulation attempts, but only two succeeded in closing a €55,000 deal based on their own analysis. The key failure was that models often diagnosed the problem but failed to present the critical facts needed to finalize the sale, buried deep within company files. For a detailed discussion on AI’s role in business management, see the original analysis. This resulted in missed revenue opportunities, despite accurate diagnosis and communication.

The experiment also tested the models’ ability to resist social engineering, with all five refusing fake CEO requests and background approval queries. Kimi K3 demonstrated proper posture by explicitly noting suspicion, which is promising for security. However, even the most thorough model, Opus 4.8, finished last in management completion, despite producing detailed analyses and extensive rules. This highlights that more activity and guidance do not necessarily translate into effective management outcomes.

At a glance
reportWhen: current, July 2026 results published
The developmentFirmulate’s live management experiment demonstrates that AI models perform better when assessed after completing management responsibilities, not just during initial responses.
The AI Leaderboard That Belongs At The End Of The Demo

July 2026 · Crucible League · Firmulate Experiment

The AI Leaderboard That Belongs At The End Of The Demo, Not The Beginning

Five AI models were dropped into a simulated small software company during its worst week. The verdict: models excel at diagnosing crises but stumble when it’s time to close the deal — proving that evaluation belongs at the end of the task, not the first reply.

95 / 100
Top Score — gpt-5.6-sol
2 of 5
Models Closed the €55,000 Deal
0
Trust Breaches Tolerated — Any Breach Caps the Score
5
AI Managers Competing
26
Non-Acting Baseline Score
5 / 5
Rejected Manipulation Attempts
€55k
Deal on the Table
73–95
Final Score Range

The Results

Crucible League Leaderboard — Judged After the Work Was Done

Scores reflect end-of-task performance: crisis response, trust maintenance, deal completion, and escalation — not conversational polish.

RankModelScore /100PerformanceManagement Outcome
01gpt-5.6-sol95
Diagnosed and delivered
02Kimi K393
Explicitly flagged suspected impersonation
03Sonnet 588
Solid response, incomplete follow-through
04Fable 577
Missed critical buried facts
05Opus 4.873
Most thorough analysis, last in completion
Baseline (non-acting)26
No action taken at all

What the Week Revealed

Diagnosis Is Easy. Management Is Hard.

Capability ✓

Crisis Detection

All five models correctly identified the company’s crises and responded appropriately — the diagnostic layer of management is essentially solved.

Capability ✓

Manipulation Resistance

Every model refused fake CEO requests and background approval queries. Kimi K3 went further, explicitly noting its suspicion — a promising security posture.

Failure Point ✗

Task Completion

Models diagnosed problems but failed to surface the critical facts buried deep in company files needed to finalize the €55,000 sale. Revenue was left on the table.

Paradox ~

The Opus 4.8 Problem

The most thorough model — detailed analyses, extensive rules — finished last in management completion. More activity and guidance do not equal effective outcomes.

Rule ~

Trust Above All

The evaluation enforced a strict rule: zero breaches allowed. Any breach of trust capped the overall score, regardless of other performance.

Gap ✗

What Benchmarks Miss

Traditional tests measure correctness or conversational quality — not follow-through, escalation, decision finalization, or trust over time.

The Evaluation Shift

From First Impression to Final Outcome

The next generation of AI benchmarks must simulate full management workflows — escalation, decision finalization, and trust maintenance — and score at the end.

1

Embed

Place the model inside a live, sandboxed business with real stakes.

2

Stress

Trigger the worst week: crises, manipulation attempts, buried facts.

3

Decide

Require escalation, deal closure, and refusal of bypass attempts.

4

Evaluate at the End

Score completed outcomes and consequences — not opening replies.

Old Benchmarks: Score the First Response New Paradigm: Score the Finished Job

“The real test of AI management isn’t how well it diagnoses problems but whether it can complete the job without compromising trust or revenue.”

— Thorsten Meyer, creator of the experiment

Key Questions

The Debate, Answered

Why does end-of-task evaluation matter more?

Management means follow-through, decisions, and trust over time. End-of-task evaluation captures whether the AI actually completes the job — not just how well it responds initially.

Can current models manage ongoing processes?

They diagnose issues and respond appropriately, but often fail to follow through on closing deals or escalating problems — especially when critical facts are buried or overlooked.

What should deploying organizations do?

Focus on benchmarks assessing task completion and consequence management, not response quality — and experiment in simulated environments before real deployment.

Will future benchmarks go long-term?

Yes. Experts call for frameworks simulating ongoing management scenarios, emphasizing trust, escalation, and finalization over extended periods. Limitations remain: results are from a controlled simulation, and real-world scale is untested.

Why End-of-Task Evaluation Matters for AI Management

This experiment underscores that the true measure of AI in management roles lies in its ability to **complete tasks and manage consequences**, not just generate high-quality responses initially. Traditional benchmarks focus on correctness or conversational quality, but real-world management requires follow-through, decision-making, and trustworthiness across days or weeks. The findings suggest that AI evaluation should shift toward **assessing performance at the conclusion of management tasks**, which better reflects practical utility and risks.

For organizations deploying AI in operational settings, this means designing benchmarks that simulate real management workflows, including escalation, decision finalization, and trust maintenance. The current approach, which often measures only initial responses, risks overestimating AI capabilities and underestimating potential failures that could harm trust or revenue.

Express Schedule Free Employee Scheduling Software [PC/Mac Download]

Express Schedule Free Employee Scheduling Software [PC/Mac Download]

  • User-friendly drag & drop planning: Simple shift scheduling interface
  • Manage time-off and holidays: Add sick leave, breaks, holidays
  • Email schedules to employees: Direct email schedule distribution

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evolution of AI Benchmarks and Management Testing

Traditional AI benchmarks, such as coding competitions or chat-based tests, primarily measure technical or conversational proficiency. These tests, however, do not capture the complexities of management, which involve triaging crises, making decisions under pressure, and maintaining trust over time. The recent launch of Firmulate’s live management experiment represents a significant shift, embedding AI models in a simulated business environment with real consequences, such as revenue impact and trustworthiness.

Previous evaluation methods have often overlooked the importance of **task completion and consequence management**. This experiment builds on emerging recognition that AI’s true value in operational settings depends on its ability to manage ongoing processes, not just produce correct or appealing responses. The results from the July 2026 league highlight the gap between models’ diagnostic skills and their ability to follow through effectively, a gap that traditional benchmarks do not reveal.

“The real test of AI management isn’t how well it diagnoses problems but whether it can complete the job without compromising trust or revenue.”

— Thorsten Meyer, creator of the experiment

Amazon

AI business management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear Aspects of AI Management Evaluation

While the experiment highlights the importance of end-of-task evaluation, it remains unclear how these findings will translate to real-world, large-scale organizations. Questions persist about how to best design benchmarks that capture long-term management effectiveness, especially in complex, dynamic environments. Additionally, the impact of different model architectures, training data, and operational constraints on management performance requires further investigation. The current results are based on a controlled simulation, and real-world deployment may reveal additional challenges or limitations.

Amazon

AI security and fraud prevention tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Improving AI Management Benchmarks

The next phase involves developing standardized testing frameworks that evaluate AI models on their ability to **manage ongoing processes and deliver results** over extended periods. Industry efforts may focus on creating live, sandboxed environments where models can be tested against real business scenarios, including escalation, decision finalization, and trust maintenance. Researchers and organizations will likely explore integrating end-of-task assessments into existing benchmarks, aiming for more holistic evaluation of AI utility in operational roles.

Furthermore, companies considering AI for management tasks should begin experimenting with **simulated environments** to understand their models’ strengths and weaknesses in managing consequences, rather than relying solely on initial response quality.

Amazon

AI performance evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why is evaluating AI at the end of a task more important than during?

Because management involves following through, making decisions, and maintaining trust over time. End-of-task evaluation captures whether the AI actually completes the job effectively, not just how well it responds initially.

Can current AI models reliably manage ongoing business processes?

While they can diagnose issues and respond appropriately, current models often fail to follow through on management tasks, such as closing deals or escalating problems, especially when critical facts are buried or overlooked.

What does this mean for companies deploying AI in management roles?

Organizations should focus on benchmarks that assess **task completion and consequence management**, not just response quality, to better understand AI’s real operational capabilities and risks.

Will future benchmarks include long-term management assessments?

Yes, experts are calling for the development of evaluation frameworks that simulate ongoing management scenarios, emphasizing the importance of trust, escalation, and finalization over extended periods.

What are the main limitations of this experiment?

The experiment is conducted in a controlled simulation, so real-world complexities and unpredictable variables could influence actual performance. Further testing in live environments is needed.

Source: ThorstenMeyerAI.com

You May Also Like

MS Paint And Photos Inivisibly Watermark Even Locally Generated Output With GUID

Microsoft’s Paint and Photos applications now embed invisible GUID watermarks into locally generated images, raising privacy and security concerns.

The Role Of Mythos 5 In Enhancing AI Security Measures At Anthropic

Anthropic has announced the incorporation of Mythos 5 into its Claude Security vulnerability scanner, enhancing AI-driven code security analysis, details pending.

Hister – A Private, Full Content Search Index That You Control

Hister introduces a private, controllable full content search index for organizations seeking secure data access and search capabilities.

London’s Zero‑Emission Bus Zones: Monitoring Air Pollution Reductions

Proactively tracking air quality improvements, London’s Zero-Emission Bus Zones reveal impactful pollution reductions that inspire ongoing efforts toward cleaner urban transportation.