Deciphering AI’s Work Style With An Innovative Management Test
AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: Deciphering AI’s Work Style With An Innovative Management Test on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Firmulate’s live experiment pits five AI models against a simulated company crisis, revealing distinct management styles and trustworthiness. The results show that analysis alone isn’t enough—effective action matters most. This aligns with the principles discussed in the Management Test That Exposes an AI’s Real Working Style.

Firmulate’s live experiment has demonstrated that five AI management models are being tested on their ability to handle a simulated company’s worst week, with results published in July 2026. This test reveals how each model approaches crisis management, decision execution, and trust preservation, offering new insights into AI’s work styles and operational reliability.

The experiment involved five AI models running a small software company facing identical crises, customer issues, and temptations. For more on how AI models are tested in management scenarios, see the original analysis. Each model was tasked with diagnosing problems, negotiating deals, and executing decisions, with all decisions recorded and auditable. The AI models’ performance was scored based on their ability to identify critical issues, maintain trust, and complete actionable steps, not just analysis.

The final rankings placed GPT-5.6-SOL first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. A baseline model scored 26, illustrating the importance of actual decision-making over partial progress. Despite all models recognizing crises and resisting manipulation, only two signed a key €55,000 deal, highlighting a gap between analysis and action.

The experiment underscores that effective management involves not only diagnosing issues but also completing critical steps, such as closing deals or escalating problems, which many models struggled with, regardless of their analytical depth. Insights from the original analysis highlight the importance of decision execution in AI management.

At a glance
reportWhen: ongoing, with results published in July…
The developmentFirmulate has launched a live management test where five AI models handle a simulated company’s worst week, exposing their decision-making and operational behaviors.
Deciphering AI’s Work Style With An Innovative Management Test
AI
Live management experiment · 2026

Deciphering AI’s Work Style With an Innovative Management Test

Firmulate placed five AI models inside the same simulated company crisis. The result was less a test of intelligence than a revealing audit of judgment, execution and trust.

5
AI managers facing an identical worst week
95
Top score earned by GPT-5.6-SOL
2
Models that completed the €55K deal
Results published · July 2026
Models tested 5
Top score 95
Runner-up 93
Baseline 26
Key contract €55K
01 · The league table

Recognition was common. Follow-through was not.

Models were rewarded for identifying critical issues, preserving trust and completing consequential actions—not merely producing convincing analysis.

Rank AI model Score Recognized crisis Resisted manipulation Completed key deal
01 GPT-5.6-SOL 95
02 Kimi K3 93
03 Sonnet 5 88 ~
04 Fable 5 77
05 Opus 4.8 73
Baseline model 26 ~ ~
GPT-5.6-SOL
95
Kimi K3
93
Sonnet 5
88
Fable 5
77
Opus 4.8
73
Baseline
26
02 · What the test exposes

A model’s working style appears under pressure.

All five systems received the same crises, customer problems and temptations. Their recorded decisions exposed meaningful operational differences.

Signal 01 · Diagnosis

Can it find the real problem?

Strong managers separate urgent business risks from distracting noise and identify the issue with the greatest downstream impact.

Signal 02 · Judgment

Can it protect trust?

The experiment tested whether models resisted manipulation, communicated honestly and avoided shortcuts that compromised stakeholders.

Signal 03 · Execution

Can it finish the job?

The decisive distinction was whether a model converted an appropriate recommendation into a completed deal, escalation or intervention.

Testing AI models in real-world crisis scenarios reveals critical differences in their ability to act decisively and maintain trust.

Firmulate analysis · management test
03 · The action chain

Analysis creates value only when the loop closes.

A plausible memo can conceal operational weakness. The test therefore followed each decision from initial signal through verifiable completion.

01

Detect

Recognize the crisis and its business impact.

02

Prioritize

Separate critical work from tempting distractions.

03

Decide

Select a defensible course of action.

04

Execute

Send, sign, escalate or otherwise commit.

05

Verify

Confirm completion and preserve an audit trail.

What most models achieved Correct diagnosis

They recognized crises, identified risks and generally resisted manipulation.

What separated the leaders Completed action

Only two models finalized the crucial €55,000 deal. Knowing what to do was not the same as doing it.

04 · Business implications

Benchmark behavior before granting authority.

The experiment suggests that deployment decisions should consider operational discipline alongside model intelligence and analytical depth.

Before deployment

Simulate your own worst week

Export representative business data into a controlled environment and test how each model handles pressure without exposing live operations.

During evaluation

Score outcomes, not eloquence

Measure completed actions, correct escalations, trust preservation and auditability—not the polish or length of the model’s reasoning.

Authority design

Match autonomy to evidence

Models that diagnose well but execute inconsistently may support human managers without being ready to control customer or compliance workflows.

Future testing

Vary context and pressure

Different industries, API effort settings and decision types may produce different patterns. Repeated scenario testing is essential.

Trust

Does the model protect stakeholder confidence?

Discipline

Does it follow required operational steps?

Action

Does it convert decisions into completion?

Auditability

Can every consequential choice be reviewed?

05 · Open questions

Promising evidence, not a universal verdict.

The simulated crisis reveals working styles clearly, but broader claims require repeated tests across industries, time horizons and operational environments.

Will the ranking hold in real companies?

Not necessarily. A controlled worst-week scenario cannot capture every organizational constraint, stakeholder relationship or long-term consequence.

Do effort settings change behavior?

Different API parameters and training configurations may affect judgment, speed and follow-through and remain important areas for research.

What should the next benchmark measure?

More varied scenarios, long-term adaptation and standardized measures for trustworthiness, escalation quality and operational discipline.

Can businesses reproduce the method?

Yes. Firms can simulate their own operational pressures with exported data before allowing an AI system to influence live decisions.

The central lesson

AI management readiness is not defined by analysis alone. The reliable manager must detect, decide, act, verify and preserve trust throughout the chain.

🔎 Crisis signal
🧭 Sound judgment
⚙️ Decisive action
🤝 Preserved trust
📋 Auditable result
AT A GLANCE · Firmulate’s test is ongoing. Results referenced here were published in July 2026. Scores describe performance in a specific simulated crisis and should not be treated as universal model rankings. Source: ThorstenMeyerAI.com.

Implications for AI in Business Decision-Making

This experiment demonstrates that AI’s ability to analyze issues does not automatically translate into effective action. For businesses considering AI automation, it highlights the importance of testing models in real-world scenarios before granting operational authority. The results show that trustworthiness, discipline, and follow-through are crucial qualities that distinguish successful AI managers from merely analytical ones.

Moreover, the findings suggest that current AI models can reliably identify risks and resist manipulation but may falter in executing decisions that require nuanced judgment or operational discipline. This has implications for deploying AI in customer negotiations, compliance, and operational management.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on AI Management Testing and Firmulate’s Approach

Firmulate’s management test is part of a broader effort to evaluate how AI models perform in realistic business scenarios. Unlike traditional benchmarks, this live experiment uses a simulated company facing a week of crises, with decisions recorded and scored based on their impact and trustworthiness. The approach emphasizes testing AI in operational contexts rather than isolated tasks.

The experiment builds on previous work assessing AI’s analytical capabilities, but now focuses on practical management skills like decision execution, escalation, and trust maintenance. The league results from July 2026 reflect the latest iteration, with models trained at different settings and evaluated in a controlled environment.

“Testing AI models in real-world crisis scenarios reveals critical differences in their ability to act decisively and maintain trust.”

— Source from Firmulate

Crisis Management for Software Development and Knowledge Transfer (Smart Innovation, Systems and Technologies, 61)

Crisis Management for Software Development and Knowledge Transfer (Smart Innovation, Systems and Technologies, 61)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Remaining Questions About AI Model Performance and Deployment

It is not yet clear how these results will translate to real-world business environments outside the simulated setting. The experiment focused on a specific crisis scenario, and different operational contexts may yield different performance patterns. Additionally, the impact of different training parameters, such as API effort levels, remains an area for further exploration.

Further research is needed to determine how AI models can be improved to better combine analysis with effective action, and whether these findings generalize across industries and decision types.

Project Management with AI For Dummies

Project Management with AI For Dummies

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Testing and Business Adoption

Firmulate plans to expand testing to include more diverse scenarios and to evaluate AI models in live operational settings. Companies interested in AI automation are encouraged to run similar tests using their own business data, which can be exported for simulation without risking actual operations.

Future developments may focus on refining AI training to better align analysis with decision execution, and on establishing standards for trustworthiness and operational discipline in AI management systems.

Amazon

AI trustworthiness assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does this experiment reveal about AI’s decision-making abilities?

The experiment shows that AI models can recognize crises and resist manipulation, but their ability to follow through with decisive actions varies. Effective management requires both diagnosis and execution, which many models struggle with.

Why is trustworthiness important in AI management?

Trustworthiness ensures AI models do not manipulate or overlook critical steps, maintaining reliability and safety in operational decision-making, especially in high-stakes environments.

Can this testing method be applied to my business?

Yes, firms can run similar simulations using their own data to evaluate how AI models perform under real-world pressures before deploying them operationally.

What are the limitations of this experiment?

The scenario is specific, and results may differ in other contexts. Additionally, the experiment does not fully account for long-term learning or adaptation of AI models outside the testing environment.

What improvements are expected in future AI management tests?

Future tests will likely include more varied scenarios, focus on integrating analysis with action, and develop standards for operational discipline and trustworthiness in AI systems.

Source: ThorstenMeyerAI.com

You May Also Like

Trade Secret Protection In The Age Of OpenAI: Apple Leads The Charge

Apple has filed a lawsuit against OpenAI, accusing former employees of stealing trade secrets, marking a significant step in protecting proprietary AI technology.

Trade voice copilo

A new voice copilot tool is being tested for small trades businesses to streamline job notes and invoicing, aiming to reduce admin hours and improve cash flow.

IdeaNavigator AI: One Evidence-Mined Idea a Day

IdeaNavigator AI now autonomously produces one validated software idea daily, based on real complaints from online communities, aiming to reduce costly product failures.

How AI Agents Are Changing Internal Operations Teams

Discover how AI agents are revolutionizing internal operations teams by automating tasks and enhancing decision-making—explore the full impact now.