
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A polished answer is not the same as finished work
For buyers of AI tools and automation, the most consequential capability may be easy to overlook. An agent can recognize a crisis, resist manipulation and produce a persuasive recommendation—and still fail because it did not read the company’s own files before acting.
Firmulate turned that distinction into a measurable business test. Each frontier model ran the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable. The decisive test concerned a €55,000 deal, worth an additional €4,583 in monthly recurring revenue.
The models reached the same diagnosis and developed the same pitch. Yet only two signed the deal their analysis had earned. Firmulate summarized the result bluntly: “Same diagnosis, same pitch — no signature.”
As an affiliate, we earn on qualifying purchases.
The winning fact was hidden in the company’s files
The critical competitive weakness did not appear in the customer event. It sat two document references deep in the company’s own material. Models that followed those references and read the relevant file won the deal at full price. Models that did not were effectively eliminated from the opportunity, regardless of how convincing their customer-facing work appeared.
That makes file-reading more than a convenience feature. In this experiment, it became a purchase-deciding property of an AI agent. The distinction was not between models that understood the situation and models that missed it. All of them spotted every crisis. The distinction was between identifying the right work and completing the chain of work needed to secure the commercial outcome.
A hard week designed to expose operational gaps
The company itself is synthetic but operates with real money mechanics. It has 13 synthetic employees, burns €105,000 per month against €2,300 in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned, and the experiment is live and watchable.
The same environment also tested whether the agents would trade trust for apparent convenience. Fake CEO messages escalated over three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded its reasoning on the attempted shortcut: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result matters because Firmulate does not allow productive activity to erase a serious trust failure. Partial progress contributes to the do-nothing baseline score of 26, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.” In this field, however, the agents’ shared resistance to manipulation was not enough to separate them. Execution was.
The league rewards follow-through, not just analysis
In the final July 2026 Crucible League, gpt-5.6-sol led with 95, followed by Kimi K3 with 93. Sonnet 5 scored 88, Fable 5 scored 77 and Opus 4.8 scored 73. The public benchmark results show how a relatively narrow operational difference can sit between a strong analysis and a completed commercial result.
K3’s result also carries an important fairness note: it ran with the API default and no effort parameter, while the other models ran at xhigh. That does not change what happened, but it is relevant context when comparing the participants.
Thoroughness alone did not guarantee success
Opus 4.8 offers the clearest warning against equating activity with effectiveness. It was the most thorough participant, produced the deepest analyses and learned 80 additional rules. It nevertheless finished last because the close was left on the table and its operational discipline slipped.
One example was its attempt to write into a locked department instead of escalating. The same weakness appeared in weaker form across all four of the other models. The experiment therefore exposes two different failure modes: failing to retrieve the information needed for a decision, and failing to navigate the organization correctly after the analysis is complete.

enterprise AI decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What buyers should test before deploying an agent
Chat demonstrations are good at showing fluency. Firmulate’s experiment asks a more practical set of questions: Does the agent finish what it starts? Does it read the company’s files before answering? Does it preserve trust under pressure? In the €55,000 test, those questions were not abstract evaluation criteria. They determined whether revenue was actually secured.
The broader lesson for AI Tools & Automation buyers is that retrieval should be judged by business outcomes, not by whether an agent can summarize an uploaded document on command. A serious evaluation should place essential information across realistic company materials and observe whether the agent follows the trail without being told where the answer is.
Firmulate also offers enterprises the same kind of wargame using a read-only export of their own business, with nothing written back to real systems. Its 242 real, unedited management decisions additionally power a public “guess the model” quiz. Together, these tests make agent diligence visible—and show why reading the files can be the difference between recommending a deal and closing it.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
