AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

A live experiment by Firmulate tested AI models managing a small company during its worst week. Results reveal models excel at diagnosing issues but fail to complete management tasks, highlighting the need for evaluation at the end of AI interactions.

Firmulate’s latest experiment places AI models in the role of managing a small software company during its most challenging week. The results show that models can identify crises and respond appropriately but often fail to complete management tasks, such as closing deals or escalating issues. This reveals that evaluating AI performance only on initial responses misses critical aspects of real-world management, emphasizing the need to assess models at the end of their decision-making process. Insights on comprehensive AI evaluation can be found in the original analysis.

The experiment involved five AI managers competing in the July 2026 Crucible League, with gpt-5.6-sol achieving the top score of 95 out of 100, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. For more on evaluating AI models in management scenarios, see the original analysis. A baseline score of 26 was recorded for a simple, non-acting model. The evaluation emphasized trust, with a strict rule: no breaches allowed, and any breach capped the overall score.

All models correctly identified crises and rejected manipulation attempts, but only two succeeded in closing a €55,000 deal based on their own analysis. The key failure was that models often diagnosed the problem but failed to present the critical facts needed to finalize the sale, buried deep within company files. For a detailed discussion on AI’s role in business management, see the original analysis. This resulted in missed revenue opportunities, despite accurate diagnosis and communication.

The experiment also tested the models’ ability to resist social engineering, with all five refusing fake CEO requests and background approval queries. Kimi K3 demonstrated proper posture by explicitly noting suspicion, which is promising for security. However, even the most thorough model, Opus 4.8, finished last in management completion, despite producing detailed analyses and extensive rules. This highlights that more activity and guidance do not necessarily translate into effective management outcomes.

At a glance
reportWhen: current, July 2026 results published
The developmentFirmulate’s live management experiment demonstrates that AI models perform better when assessed after completing management responsibilities, not just during initial responses.

Why End-of-Task Evaluation Matters for AI Management

This experiment underscores that the true measure of AI in management roles lies in its ability to **complete tasks and manage consequences**, not just generate high-quality responses initially. Traditional benchmarks focus on correctness or conversational quality, but real-world management requires follow-through, decision-making, and trustworthiness across days or weeks. The findings suggest that AI evaluation should shift toward **assessing performance at the conclusion of management tasks**, which better reflects practical utility and risks.

For organizations deploying AI in operational settings, this means designing benchmarks that simulate real management workflows, including escalation, decision finalization, and trust maintenance. The current approach, which often measures only initial responses, risks overestimating AI capabilities and underestimating potential failures that could harm trust or revenue.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Evolution of AI Benchmarks and Management Testing

Traditional AI benchmarks, such as coding competitions or chat-based tests, primarily measure technical or conversational proficiency. These tests, however, do not capture the complexities of management, which involve triaging crises, making decisions under pressure, and maintaining trust over time. The recent launch of Firmulate’s live management experiment represents a significant shift, embedding AI models in a simulated business environment with real consequences, such as revenue impact and trustworthiness.

Previous evaluation methods have often overlooked the importance of **task completion and consequence management**. This experiment builds on emerging recognition that AI’s true value in operational settings depends on its ability to manage ongoing processes, not just produce correct or appealing responses. The results from the July 2026 league highlight the gap between models’ diagnostic skills and their ability to follow through effectively, a gap that traditional benchmarks do not reveal.

“The real test of AI management isn’t how well it diagnoses problems but whether it can complete the job without compromising trust or revenue.”

— Thorsten Meyer, creator of the experiment

Amazon

AI project management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear Aspects of AI Management Evaluation

While the experiment highlights the importance of end-of-task evaluation, it remains unclear how these findings will translate to real-world, large-scale organizations. Questions persist about how to best design benchmarks that capture long-term management effectiveness, especially in complex, dynamic environments. Additionally, the impact of different model architectures, training data, and operational constraints on management performance requires further investigation. The current results are based on a controlled simulation, and real-world deployment may reveal additional challenges or limitations.

Amazon

AI decision-making evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Improving AI Management Benchmarks

The next phase involves developing standardized testing frameworks that evaluate AI models on their ability to **manage ongoing processes and deliver results** over extended periods. Industry efforts may focus on creating live, sandboxed environments where models can be tested against real business scenarios, including escalation, decision finalization, and trust maintenance. Researchers and organizations will likely explore integrating end-of-task assessments into existing benchmarks, aiming for more holistic evaluation of AI utility in operational roles.

Furthermore, companies considering AI for management tasks should begin experimenting with **simulated environments** to understand their models’ strengths and weaknesses in managing consequences, rather than relying solely on initial response quality.

Amazon

AI security and social engineering resistance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

Why is evaluating AI at the end of a task more important than during?

Because management involves following through, making decisions, and maintaining trust over time. End-of-task evaluation captures whether the AI actually completes the job effectively, not just how well it responds initially.

Can current AI models reliably manage ongoing business processes?

While they can diagnose issues and respond appropriately, current models often fail to follow through on management tasks, such as closing deals or escalating problems, especially when critical facts are buried or overlooked.

What does this mean for companies deploying AI in management roles?

Organizations should focus on benchmarks that assess **task completion and consequence management**, not just response quality, to better understand AI’s real operational capabilities and risks.

Will future benchmarks include long-term management assessments?

Yes, experts are calling for the development of evaluation frameworks that simulate ongoing management scenarios, emphasizing the importance of trust, escalation, and finalization over extended periods.

What are the main limitations of this experiment?

The experiment is conducted in a controlled simulation, so real-world complexities and unpredictable variables could influence actual performance. Further testing in live environments is needed.

Source: ThorstenMeyerAI.com

You May Also Like

Halcyon Video – A 3D Video Store For Your Media Server

Halcyon Video introduces a new 3D video store designed for media servers, expanding media options with immersive content. Details are confirmed for early access.

GrapheneOS’ Rewritten Messages App Is Released

GrapheneOS has launched a new, rewritten Messages app aimed at enhancing privacy and security for users of its privacy-focused OS.

Oslo’s Electric Bus Adoption: Policy Support and Lessons Learned

Keen insights into Oslo’s electric bus success reveal strategies that could transform urban transit worldwide—discover the lessons learned and their future implications.

California’s Innovative Clean Transit Rule: Tracking Agency Compliance to 2040

Find out how California’s Clean Transit Rule tracks agency progress toward 2040’s zero-emission goals and what strategies ensure success.