📊 Full opportunity report: The Final Word In AI: The Leaderboard After The Demo Ends on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
The Firmulate experiment concluded with GPT-5.6-SOL leading the leaderboard, highlighting that management skills and trustworthiness matter more than just technical responses. The results challenge traditional AI benchmarks by focusing on real-world decision management.
The Firmulate live experiment concluded in July 2026, with GPT-5.6-SOL ranking first among five AI models, demonstrating that management performance and trustworthiness are vital metrics for AI in business contexts. This marks a significant shift from traditional benchmarks focused solely on technical output or conversational preference, emphasizing the importance of decision-making under real-world conditions.
The experiment involved five AI models managing a simulated small software company facing multiple crises over a week, as detailed in the original analysis. The models were evaluated on their ability to diagnose issues, communicate effectively, and, critically, maintain trust by avoiding breaches. GPT-5.6-SOL scored 95 points, narrowly outperforming Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. A baseline model scored 26, showing partial progress.
Unlike typical chat benchmarks, this experiment enforced strict trust standards: a single breach of trust capped the score, reflecting real-world consequences. All models identified crises and rejected manipulation attempts; however, only two successfully signed a €55,000 deal after diagnosing the company’s issues. The key failure was not in diagnosis but in execution—models that retrieved the correct information from internal files were more successful in closing deals, highlighting the importance of accurate data retrieval over superficial responses.
Additionally, models demonstrated robustness against social engineering attempts, refusing fake CEO messages and impersonation tricks. Yet, even the best models struggled with completing managerial tasks effectively, such as escalating issues or following through on decisions. Opus 4.8, despite detailed analysis and extensive rule application, finished last due to lapses in discipline and escalation, illustrating that effort and thoroughness do not necessarily translate into effective management outcomes.
Why Management Skills and Trust Matter in AI Evaluation
The results underscore that traditional AI benchmarks—focused on technical accuracy or conversational fluency—are insufficient for real-world business applications. The experiment demonstrates that management quality, trustworthiness, and decision execution are critical factors, shaping how AI can be integrated into organizational workflows. This shift could influence future AI development and enterprise adoption strategies, emphasizing holistic evaluation over isolated performance metrics.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background of AI Benchmarks and the Firmulate Experiment
Historically, AI benchmarks have prioritized coding accuracy, language fluency, or user preference, often measured through controlled tests or chat arenas. These metrics, however, do not capture how models perform in managing complex, dynamic situations involving multiple stakeholders, trust, and consequences. The Firmulate experiment, launched as a live test in July 2026, introduces a new paradigm by simulating a small company’s week of crises, requiring models to diagnose, communicate, escalate, and close deals under real-world pressures. This approach exposes the gap between technical prowess and practical management capabilities.
Prior to this, AI evaluation largely ignored the management aspect, focusing instead on isolated tasks. The Firmulate experiment’s design—versioned decision-making, auditable actions, and real financial consequences—aims to measure how models handle organizational complexity and trust, providing a more comprehensive assessment of AI readiness for enterprise deployment.
“This experiment shows that management quality, not just chat quality, should be its own category in AI evaluation.”
— Thorsten Meyer, lead researcher at Firmulate
AI trustworthiness evaluation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Remaining Questions About Model Performance and Application
While the leaderboard ranks models based on their performance in this simulated environment, it remains unclear how these results translate to real-world business operations across different industries. The experiment’s scope was limited to a specific scenario involving a small software company, and broader validation is needed to confirm whether these findings hold in more complex or diverse organizational contexts. Additionally, the long-term reliability of models in sustained management roles, especially under evolving crises, is still untested.
It is also not yet clear how future models will adapt to the management challenges highlighted, or whether new evaluation standards will be adopted industry-wide. The impact of different training data, fine-tuning, and integration strategies on management performance remains an open question.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Evaluation and Adoption
Following the experiment, industry stakeholders are likely to reconsider how they evaluate AI models for enterprise use, emphasizing trust, decision-making, and management capabilities. Companies may run their own wargames or simulations, similar to Firmulate’s, to test models before deployment. Developers are expected to focus on improving models’ ability to retrieve accurate information, escalate appropriately, and maintain trust under pressure.
Further research and larger-scale experiments are anticipated to refine these benchmarks, potentially leading to new standards that prioritize management effectiveness. Regulatory and ethical considerations around trustworthiness and decision accountability will also shape the future landscape of AI in organizations.
Finally, real-world deployment will require ongoing monitoring, with organizations implementing continuous testing to ensure models uphold management quality and trustworthiness over time, especially as models evolve or are fine-tuned for specific tasks.

AI for Public Relations: A How-To Guide for Implementation and Management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does the leaderboard tell us about AI’s management capabilities?
The leaderboard indicates that some models, like GPT-5.6-SOL, can perform well in diagnosing issues, communicating effectively, and maintaining trust—key aspects of management. However, technical accuracy alone does not guarantee success in real-world management tasks, which require execution, escalation, and discipline.
Why is trustworthiness emphasized in the experiment?
Trustworthiness is critical because a breach can have immediate negative consequences, such as losing deals or damaging reputation. The experiment capped scores after a trust breach, highlighting that reliability and ethical behavior are essential for AI models managing organizational decisions.
Can these results be applied to different industries or larger organizations?
While the experiment offers valuable insights, its scope was limited to a small software company scenario. Broader validation across industries and organizational sizes is needed to confirm applicability, and ongoing research will clarify how well these benchmarks generalize.
What should companies consider before deploying AI for management tasks?
Organizations should evaluate whether models can read and interpret organizational files accurately, escalate issues appropriately, and maintain trust under pressure. Running simulations or wargames can help assess these capabilities before full deployment.
What is the significance of the experiment’s findings for AI development?
The findings suggest that future AI development should prioritize management skills, trustworthiness, and decision execution over mere conversational or technical performance. This shift could redefine standards for enterprise AI adoption.
Source: ThorstenMeyerAI.com