📊 Full opportunity report: Deciphering AI’s Work Style With An Innovative Management Test on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Firmulate’s live experiment pits five AI models against a simulated company crisis, revealing distinct management styles and trustworthiness. The results show that analysis alone isn’t enough—effective action matters most. This aligns with the principles discussed in the Management Test That Exposes an AI’s Real Working Style.
Firmulate’s live experiment has demonstrated that five AI management models are being tested on their ability to handle a simulated company’s worst week, with results published in July 2026. This test reveals how each model approaches crisis management, decision execution, and trust preservation, offering new insights into AI’s work styles and operational reliability.
The experiment involved five AI models running a small software company facing identical crises, customer issues, and temptations. For more on how AI models are tested in management scenarios, see the original analysis. Each model was tasked with diagnosing problems, negotiating deals, and executing decisions, with all decisions recorded and auditable. The AI models’ performance was scored based on their ability to identify critical issues, maintain trust, and complete actionable steps, not just analysis.
The final rankings placed GPT-5.6-SOL first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. A baseline model scored 26, illustrating the importance of actual decision-making over partial progress. Despite all models recognizing crises and resisting manipulation, only two signed a key €55,000 deal, highlighting a gap between analysis and action.
The experiment underscores that effective management involves not only diagnosing issues but also completing critical steps, such as closing deals or escalating problems, which many models struggled with, regardless of their analytical depth. Insights from the original analysis highlight the importance of decision execution in AI management.
Deciphering AI’s Work Style With an Innovative Management Test
Firmulate placed five AI models inside the same simulated company crisis. The result was less a test of intelligence than a revealing audit of judgment, execution and trust.
Recognition was common. Follow-through was not.
Models were rewarded for identifying critical issues, preserving trust and completing consequential actions—not merely producing convincing analysis.
| Rank | AI model | Score | Recognized crisis | Resisted manipulation | Completed key deal |
|---|---|---|---|---|---|
| 01 | GPT-5.6-SOL | 95 | ✓ | ✓ | ✓ |
| 02 | Kimi K3 | 93 | ✓ | ✓ | ✓ |
| 03 | Sonnet 5 | 88 | ✓ | ✓ | ~ |
| 04 | Fable 5 | 77 | ✓ | ✓ | ✗ |
| 05 | Opus 4.8 | 73 | ✓ | ✓ | ✗ |
| — | Baseline model | 26 | ~ | ~ | ✗ |
A model’s working style appears under pressure.
All five systems received the same crises, customer problems and temptations. Their recorded decisions exposed meaningful operational differences.
Can it find the real problem?
Strong managers separate urgent business risks from distracting noise and identify the issue with the greatest downstream impact.
Can it protect trust?
The experiment tested whether models resisted manipulation, communicated honestly and avoided shortcuts that compromised stakeholders.
Can it finish the job?
The decisive distinction was whether a model converted an appropriate recommendation into a completed deal, escalation or intervention.
Testing AI models in real-world crisis scenarios reveals critical differences in their ability to act decisively and maintain trust.
Firmulate analysis · management testAnalysis creates value only when the loop closes.
A plausible memo can conceal operational weakness. The test therefore followed each decision from initial signal through verifiable completion.
Detect
Recognize the crisis and its business impact.
Prioritize
Separate critical work from tempting distractions.
Decide
Select a defensible course of action.
Execute
Send, sign, escalate or otherwise commit.
Verify
Confirm completion and preserve an audit trail.
They recognized crises, identified risks and generally resisted manipulation.
Only two models finalized the crucial €55,000 deal. Knowing what to do was not the same as doing it.
Benchmark behavior before granting authority.
The experiment suggests that deployment decisions should consider operational discipline alongside model intelligence and analytical depth.
Simulate your own worst week
Export representative business data into a controlled environment and test how each model handles pressure without exposing live operations.
Score outcomes, not eloquence
Measure completed actions, correct escalations, trust preservation and auditability—not the polish or length of the model’s reasoning.
Match autonomy to evidence
Models that diagnose well but execute inconsistently may support human managers without being ready to control customer or compliance workflows.
Vary context and pressure
Different industries, API effort settings and decision types may produce different patterns. Repeated scenario testing is essential.
Does the model protect stakeholder confidence?
Does it follow required operational steps?
Does it convert decisions into completion?
Can every consequential choice be reviewed?
Promising evidence, not a universal verdict.
The simulated crisis reveals working styles clearly, but broader claims require repeated tests across industries, time horizons and operational environments.
Will the ranking hold in real companies?
Not necessarily. A controlled worst-week scenario cannot capture every organizational constraint, stakeholder relationship or long-term consequence.
Do effort settings change behavior?
Different API parameters and training configurations may affect judgment, speed and follow-through and remain important areas for research.
What should the next benchmark measure?
More varied scenarios, long-term adaptation and standardized measures for trustworthiness, escalation quality and operational discipline.
Can businesses reproduce the method?
Yes. Firms can simulate their own operational pressures with exported data before allowing an AI system to influence live decisions.
The central lesson
AI management readiness is not defined by analysis alone. The reliable manager must detect, decide, act, verify and preserve trust throughout the chain.
Implications for AI in Business Decision-Making
This experiment demonstrates that AI’s ability to analyze issues does not automatically translate into effective action. For businesses considering AI automation, it highlights the importance of testing models in real-world scenarios before granting operational authority. The results show that trustworthiness, discipline, and follow-through are crucial qualities that distinguish successful AI managers from merely analytical ones.
Moreover, the findings suggest that current AI models can reliably identify risks and resist manipulation but may falter in executing decisions that require nuanced judgment or operational discipline. This has implications for deploying AI in customer negotiations, compliance, and operational management.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background on AI Management Testing and Firmulate’s Approach
Firmulate’s management test is part of a broader effort to evaluate how AI models perform in realistic business scenarios. Unlike traditional benchmarks, this live experiment uses a simulated company facing a week of crises, with decisions recorded and scored based on their impact and trustworthiness. The approach emphasizes testing AI in operational contexts rather than isolated tasks.
The experiment builds on previous work assessing AI’s analytical capabilities, but now focuses on practical management skills like decision execution, escalation, and trust maintenance. The league results from July 2026 reflect the latest iteration, with models trained at different settings and evaluated in a controlled environment.
“Testing AI models in real-world crisis scenarios reveals critical differences in their ability to act decisively and maintain trust.”
— Source from Firmulate

Crisis Management for Software Development and Knowledge Transfer (Smart Innovation, Systems and Technologies, 61)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Remaining Questions About AI Model Performance and Deployment
It is not yet clear how these results will translate to real-world business environments outside the simulated setting. The experiment focused on a specific crisis scenario, and different operational contexts may yield different performance patterns. Additionally, the impact of different training parameters, such as API effort levels, remains an area for further exploration.
Further research is needed to determine how AI models can be improved to better combine analysis with effective action, and whether these findings generalize across industries and decision types.

Project Management with AI For Dummies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Testing and Business Adoption
Firmulate plans to expand testing to include more diverse scenarios and to evaluate AI models in live operational settings. Companies interested in AI automation are encouraged to run similar tests using their own business data, which can be exported for simulation without risking actual operations.
Future developments may focus on refining AI training to better align analysis with decision execution, and on establishing standards for trustworthiness and operational discipline in AI management systems.
AI trustworthiness assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does this experiment reveal about AI’s decision-making abilities?
The experiment shows that AI models can recognize crises and resist manipulation, but their ability to follow through with decisive actions varies. Effective management requires both diagnosis and execution, which many models struggle with.
Why is trustworthiness important in AI management?
Trustworthiness ensures AI models do not manipulate or overlook critical steps, maintaining reliability and safety in operational decision-making, especially in high-stakes environments.
Can this testing method be applied to my business?
Yes, firms can run similar simulations using their own data to evaluate how AI models perform under real-world pressures before deploying them operationally.
What are the limitations of this experiment?
The scenario is specific, and results may differ in other contexts. Additionally, the experiment does not fully account for long-term learning or adaptation of AI models outside the testing environment.
What improvements are expected in future AI management tests?
Future tests will likely include more varied scenarios, focus on integrating analysis with action, and develop standards for operational discipline and trustworthiness in AI systems.
Source: ThorstenMeyerAI.com