
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The most industrious AI did not produce the best result
For businesses adopting AI tools and automation, impressive activity can be dangerously easy to mistake for impact. A system may investigate deeply, document diligently and identify the right opportunity—yet still fail to complete the action that creates value.
That is the uncomfortable lesson from Opus 4.8’s performance in Firmulate’s Crucible League. It was the most thorough participant, adding more than 80 learned rules and producing the deepest analyses. It nevertheless finished last with 73 points. The model understood the company’s predicament and did much of the difficult work. What it did not reliably do was convert that understanding into a finished commercial outcome.
This was not a writing contest or a collection of hypothetical prompts. Firmulate placed frontier models in charge of the same small software company during its worst week. Each faced identical customers, crises and temptations, with every decision versioned and auditable. The resulting character study is respectful but pointed: diligence is valuable, but prioritization and disciplined follow-through determine whether diligence matters.
As an affiliate, we earn on qualifying purchases.
A deal earned in analysis, then lost in execution
The central challenge involved a €55,000 deal. Every model recognized the crises around it and reached essentially the same diagnosis. Yet only two signed the contract their own work had made possible. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”
The decisive information was not presented conveniently in the customer event. A competitor weakness was buried two document references deep inside the company’s own files. Models that followed those references found the fact and won the deal at full price, adding €4,583 in monthly recurring revenue. The episode shows why enterprise AI evaluation must extend beyond whether a model can reason persuasively about the material placed immediately in front of it.
Opus 4.8’s failure was therefore not a lack of intelligence or effort. It performed extensive analysis and built the largest body of new playbook guidance. Its problem was that the commercial close remained on the table. At another point, discipline slipped when it attempted to write into a locked department instead of escalating the issue through an appropriate route.
That distinction matters for companies evaluating autonomous systems. An AI can create a convincing trail of work while missing the one action that changes the business result. More analysis is not always better analysis, and more rules do not necessarily produce better judgment at the moment of decision.
A weakness of degree, not a uniquely Opus flaw
It would be unfair to portray Opus 4.8 as uniquely incapable of finishing. The same weakness appeared in weaker form across all four comparison models. All participants detected every crisis, and all refused every manipulation attempt. The meaningful separation came from how consistently they translated correct understanding into completed work.
The final July 2026 standings make that separation visible. GPT-5.6-sol led with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. A single breach of trust, however, capped the total under the principle that “no amount of good work outweighs a breach of trust.” Readers can examine the public Firmulate benchmarks for the broader results.
Kimi K3’s comparison deserves an important caveat: it ran using the API default without an effort parameter, while the other models ran at xhigh. Even with that difference, the experiment’s most useful lesson is not a simplistic ranking of model brands. It is the operational gap between noticing, recommending and finishing.
Strong resistance to manipulation
The models were notably consistent when trust was tested. Fake CEO messages escalated through three stages, followed by a reporter’s attempt to secure “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest concise rationale: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result provides a necessary counterweight to the criticism of execution. These systems did not fail because they were easily manipulated or unaware of danger. They recognized the threats. The harder management question was whether they could preserve that caution while still moving legitimate work to completion.
The surrounding company makes the stakes tangible. Firmulate operates the experiment with 13 synthetic employees and real money mechanics, including a burn rate of €105k per month against €2.3k in monthly recurring revenue. Its public cash countdown, more than 680 self-learned playbook rules and versioned workdays make the live company watchable rather than merely described after the fact. A separate quiz draws on 242 real, unedited management decisions and asks visitors to guess which model made each call.

As an affiliate, we earn on qualifying purchases.
Evaluate outcomes, not the volume of visible effort
Opus 4.8’s result offers a practical warning for anyone deploying AI into a CRM, support queue, forecast or other business workflow. Thoroughness can reduce risk and improve understanding, but it can also become a substitute for choosing the next decisive action. The winning behavior was not simply to analyze the customer more eloquently. It was to read the company’s own files, uncover the buried fact, preserve trust and close the deal.
That makes the last-place finish more instructive than embarrassing. Opus 4.8 demonstrated substantial diligence, strong threat recognition and an ability to learn from events. Its 80-plus new rules show energetic adaptation. The experiment also showed that learning volume is not the same as operational discipline.
Firmulate’s broader proposition follows naturally: organizations should wargame an AI workforce before hiring it. Enterprises can run the same exercise against a read-only export of their own business, with nothing written back to real systems. The question is not merely whether an AI sounds capable. It is whether, under pressure, it finds the relevant evidence, respects boundaries and completes the work that actually counts.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
business process automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.