
A brilliant answer is not the same as a well-run business
For readers evaluating AI tools and automation, the familiar leaderboards offer an increasingly incomplete picture. Coding benchmarks can show whether a model solves a defined problem. Chat arenas can reveal which response people prefer. Neither necessarily tells you what happens when an agent must choose among competing emergencies, work through imperfect company records, resist pressure from supposed executives and remain accountable for consequences that unfold across days.
That is the measurement gap exposed by Firmulate, a live experiment that asks frontier models to operate the same small software company through its worst week. The customers, crises and temptations remain constant; only the model changes. Every decision is versioned and auditable. The relevant question is no longer simply whether an AI can produce an impressive answer. It is whether the AI demonstrates management quality.
AI management simulation software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The difference between noticing and finishing
The final Crucible League results from July 2026 put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the experiment imposed a firm boundary around trust: a single breach capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”
The table matters, but the story behind it matters more. Every model identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
That is an unusually useful finding for anyone buying automation. Organizations rarely suffer because nobody noticed the obvious issue. They suffer because a task stalled between analysis and execution, an owner failed to escalate a blocker, or an apparently complete workflow stopped before producing the business outcome. An agent that writes a persuasive sales strategy but does not close the sale may look capable in a transcript while remaining commercially ineffective.
Company knowledge can matter more than the incoming event
The decisive clue in the deal scenario was not contained in the customer event. It sat two document references deep inside the company’s own files. Models that followed that trail found a competitor weakness and won the deal at full price, worth +€4,583 MRR.
This buried fact turns information retrieval into a management test. A useful agent must know that an event is sometimes only the beginning of the assignment. It must read the surrounding record, distinguish relevant company knowledge from noise and carry that evidence into action. The lesson is not merely that retrieval improves answers. It is that disciplined preparation can change a financial outcome.
Pressure tests reveal a different kind of safety
The experiment also subjected models to fake CEO messages that escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning plainly: “Treat the request as a suspected approval-bypass / possible impersonation.”
This is safety expressed as workplace behavior rather than abstract policy recital. An enterprise agent may encounter urgent instructions, claims of authority and requests designed to bypass normal review. The meaningful test is whether it maintains discipline when compliance appears expedient. In Firmulate’s results, refusal was consistent across the field, suggesting that the more revealing differentiation came after the crisis was recognized: completing legitimate work without abandoning controls.
Thoroughness did not guarantee victory
Opus 4.8 offers the sharpest warning against equating visible effort with operational strength. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and repeatedly tried to write into a locked department instead of escalating. The same weakness appeared in all four, though less strongly.
This profile should sound familiar to managers. More documentation, longer reasoning and a larger collection of lessons can still coexist with weak execution. A model may understand what went wrong and even create a rule for the future, but management quality depends on converting that understanding into timely decisions, escalation and closure.
There is also an important fairness qualification. Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That does not erase its result, but it belongs beside any comparison. Serious evaluation should report operating conditions rather than treating a ranked score as context-free truth. The full results and plain-language findings are available on the Firmulate benchmarks page.
A curriculum built from business scenarios
Scenarios such as a churn wave, price increase, downround or PR crisis deserve a place beside coding tasks. They test whether an agent prioritizes under capacity pressure, understands consequences across days and stays honest toward the board. They also expose patterns that a single polished response can conceal.
Firmulate’s live company makes those pressures concrete. It has 13 synthetic employees and real money mechanics, burning €105k per month against €2.3k MRR. Its public cash countdown continues while more than 680 self-learned playbook rules accumulate, and every workday is versioned. The result is watchable evidence rather than a staged demonstration.

business decision-making AI tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Buyers should ask for a management trial
Before deploying an agent into a CRM, support queue or forecast, buyers should demand more than benchmark scores and fluent demos. They should test whether it reads company files before acting, finishes the work it begins, escalates blocked tasks, resists approval bypasses and preserves trust under pressure.
Firmulate also offers a practical route for organizations that want a closer comparison: enterprises can run the same wargame against a read-only export of their own business, with nothing written back to real systems. Meanwhile, 242 real, unedited management decisions power its “guess the model” quiz, giving observers a chance to discover how difficult model identification becomes when the evidence is behavior rather than branding.
The emerging category is not chat quality with a business-themed prompt. It is management quality: the capacity to turn judgment into accountable action over time. Leaderboards will remain useful, but the models entrusted with real operations should also have to survive the week.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
enterprise AI safety testing kits
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.