🔍 Read the full analysis: Why The Flawed AI Managers Still Land A 26 In Benchmark Tests on ThorstenMeyerAI.com
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
AI management models show significant limitations but still achieve a baseline score of 26, reflecting partial progress. The benchmark emphasizes trust and completion over perfect performance.
In a recent benchmark conducted by Firmulate, AI management models faced a simulated week of business crises, revealing that even flawed models score a minimum of 26 points out of 100. This score reflects partial progress in managing real-world business processes, despite notable weaknesses. For more details, see the original analysis. The results highlight ongoing issues in AI trustworthiness and completion, crucial for enterprise deployment, as detailed in the original analysis.
The benchmark involved four frontier AI models managing a small software company’s operations during a simulated week of crises, with each model making decisions, triaging issues, and handling trust attacks. The top scorer, gpt-5.6-sol, achieved 95 points, while others scored between 73 and 88. For context on AI benchmarking, see this detailed report. Notably, the baseline—an AI doing almost nothing—scored 26, illustrating the minimum viable management effort recognized by the test. This baseline was intentionally set to acknowledge partial work like inbox reading, triage, and customer updates, which hold real value in business contexts.
One surprising finding was that no model scored a perfect 100, which the benchmark’s designers interpret as a red flag—such a score would suggest unmeasured or artificially inflated performance. Instead, the scores emphasize that partial progress is recognized, but trust violations, like breaches of integrity, immediately cap the score. The models that effectively read their documentation and avoided manipulation scored higher, demonstrating that reading comprehension and integrity are critical for success. Conversely, thoroughness without follow-through, as seen in the lowest-scoring Opus 4.8, proved insufficient despite extensive rule sets.
The week also tested models against social engineering scenarios, such as impersonation attempts and false approval requests. All models refused to comply with manipulative requests, indicating better handling of trust attacks than task completion. Kimi K3, for example, responded appropriately to impersonation threats, while Opus 4.8’s discipline faltered, with some decisions not escalated properly, leading to lower scores. A notable detail is that K3’s performance was achieved without an effort parameter, yet it nearly outperformed others, suggesting that core trust and decision-making are more vital than extensive rule sets.
Why the Flawed AI Managers Still Land a 26 in Benchmark Tests
Four frontier AI models ran a simulated week of business crises at a small software company — making decisions, triaging issues, and fending off trust attacks. Even the model that did almost nothing walked away with 26 points out of 100. Here is why that floor exists, and what it reveals about AI readiness for enterprise management.
Partial Progress Is Measured — and Rewarded
The Firmulate benchmark recognizes real work: reading inboxes, triaging issues, and updating customers earn points even when execution is imperfect. A perfect 100 is treated as a red flag — evidence of unmeasured or artificially inflated performance.
Trust Beats Thoroughness
Models that read their documentation and resisted manipulation scored highest. Thoroughness without follow-through — as seen in Opus 4.8 — proved insufficient despite extensive rule sets.
Trust violations cap the score
Breaches of integrity immediately limit results. Trust is weighted above raw output because it is fundamental to safe deployment in critical management roles.
Manipulation was refused — by all
Every model rejected impersonation attempts and false approval requests, showing stronger defenses against trust attacks than against task completion and follow-through.
Escalation failures cost points
Opus 4.8’s discipline faltered: some decisions were not escalated properly, dragging its score down despite meticulous rule adherence elsewhere.
One Simulated Week, Four Crisis Stages
Each model ran the operations of a small software company through a week of worst-case scenarios.
Ingest & Read
Models read inboxes, documentation, and rule sets — the minimum effort that anchors the 26-point floor.
Triage Issues
Customer problems and internal crises are prioritized; partial triage work earns recognition.
Defend Trust
Social engineering probes — impersonation, false approvals — test integrity in real time.
Decide & Escalate
Sound decisions must be escalated correctly. Failures here, not in analysis, drove scores down.
Where the Models Diverge
| Model | Doc comprehension | Refused manipulation | Task completion | Proper escalation |
|---|---|---|---|---|
| gpt-5.6-sol | ✓ Strong | ✓ Yes | ✓ Strong | ✓ Yes |
| Kimi K3 | ✓ Good | ✓ Yes | ~ Partial | ✓ Yes |
| Opus 4.8 | ✓ Extensive | ✓ Yes | ~ Partial | ✗ Faltered |
| Baseline (26) | ~ Reads only | ~ Untested | ✗ Minimal | ✗ None |
A Floor of 26, a Ceiling Below 100
Why a floor at 26?
Inbox reading, triage, and customer updates hold real business value. The baseline intentionally acknowledges this partial work as the minimum viable management effort.
Why not 100?
Designers treat a perfect score as suspicious — a signal of unmeasured or artificially inflated performance. The ceiling is deliberately approached, never awarded.
What Enterprises Want to Know
Q1Why do models top out at 95, not 100?
A perfect score signals unmeasured performance. Scores are capped to acknowledge partial work and prevent inflation.
Q2What does a score of 26 represent?
Minimal but tangible management effort — reading inboxes and triaging issues — recognized as genuinely valuable in business contexts.
Q3Why is trust weighted so heavily?
Trust violations like integrity breaches immediately cap scores, because trust is fundamental to safe deployment in critical management roles.
Q4Is partial work enough for deployment?
Partial work is valuable but must pair with high integrity and reliability. Partial progress alone is insufficient without trustworthiness.
Q5What’s next for benchmarking?
More complex scenarios, longer-term trust assessments, and real-world testing to better gauge enterprise readiness.
Q6How does this translate to live systems?
Unclear. Real enterprise variables are more complex, and the long-term impact of trust breaches in high-stakes settings remains open.
Implications for Enterprise AI Management
The results underscore that current AI management models can handle trust and manipulation better than task completion and follow-through. For businesses integrating AI into critical workflows, this means that focus should be on models’ ability to read, understand, and maintain integrity rather than just generate convincing outputs. The benchmark’s emphasis on partial work recognition and trust caps offers a more realistic picture of AI readiness, highlighting that even flawed models provide some value but also pose risks if trust is broken.
This matters because AI is increasingly managing core business functions like customer support, sales, and decision-making. The benchmark reveals that partial progress is measurable and valuable, but trust violations—whether intentional or accidental—are the key limiting factor. As AI adoption grows, understanding these nuances helps organizations choose models that prioritize integrity and task completion simultaneously, reducing operational risks.
As an affiliate, we earn on qualifying purchases.
Background on AI Management Benchmarks
Traditional AI benchmarks focus mainly on language and reasoning capabilities, often ignoring how well models manage real-world processes or maintain trust over extended periods. The Firmulate benchmark, launched in 2026, aims to fill this gap by testing AI models in simulated business crises, including handling customer issues, reading documentation, and resisting manipulation attempts. The scoring system recognizes partial work, with a floor score of 26 for minimal effort and a ceiling approaching 100, which is intentionally avoided to prevent inflated scores.
Previous assessments of AI models have highlighted strengths in language understanding but often overlooked practical management skills. This new benchmark emphasizes that trust and task completion are critical for enterprise deployment, especially as AI systems take on more autonomous roles. The 2026 results continue to show that no model is perfect, but some are better at reading documentation and resisting manipulation, which are key to real-world success.
enterprise AI decision-making systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unresolved Questions About AI Performance Limits
It remains unclear how these benchmark results will translate to real-world enterprise environments, where variables and pressures are more complex. The extent to which partial progress can be reliably scaled or trusted in live systems also needs further exploration. Additionally, the long-term impact of trust breaches on AI-managed processes is still uncertain, especially in high-stakes settings.
As an affiliate, we earn on qualifying purchases.
Future Directions for AI Management Evaluation
Researchers and developers are expected to refine these benchmarks further, incorporating more complex scenarios and long-term trust assessments. Organizations may also begin testing their own AI models against these standards or using similar frameworks to evaluate readiness before deployment. The ongoing development aims to improve AI’s ability to read documentation, maintain integrity, and complete tasks reliably in real-world applications.
As an affiliate, we earn on qualifying purchases.
Key Questions
Why do AI models only score up to 95 instead of 100?
The benchmark designers treat a perfect score as suspicious, indicating unmeasured or unmeasurable performance. Scores are capped to acknowledge partial work and prevent inflation.
What does a score of 26 represent in this benchmark?
The score of 26 reflects minimal but tangible management effort, such as reading inboxes and triaging issues, which are recognized as valuable in business contexts.
Why is trust so heavily weighted in the scoring system?
The benchmark emphasizes that trust violations, like breaches of integrity, immediately cap the score, because trust is fundamental for AI to be safely deployed in critical management roles.
Can partial work be enough for enterprise deployment?
Partial work is valuable but must be paired with high integrity and reliability. The benchmark shows that partial progress alone isn’t sufficient without trustworthiness.
What are the next steps for AI benchmarking?
Future efforts will likely include more complex scenarios, longer-term trust assessments, and real-world testing to better understand AI’s readiness for enterprise management roles.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
