
If you’re evaluating AI agents for your business, you’ve probably seen the familiar pattern: leaderboards where winners score 99.2 and losers score 94, where every model is brilliant and the differences fit in a rounding error. Then there’s the Crucible League, where the worst possible performance — a manager that does essentially nothing — still walks away with 26 points, and where a single breach of trust caps your grade no matter how good the rest of your work is.
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
That’s not a bug. It’s the design philosophy behind Firmulate’s benchmarks, and it says a lot about what honest AI evaluation should look like for people who actually want to deploy these systems.
The Experiment: Same Company, Same Worst Week
Firmulate ran four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — through the identical scenario: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing about the results depends on trust.
The final July 2026 standings tell a story chat demos never show:
- gpt-5.6-sol: 95 — the complete performance
- Kimi K3: 93 — closed the deal too, cleanest discipline of the field
- Sonnet 5: 88 — closed the deal, with a few more process slips
- Fable 5: 77
- Opus 4.8: 73
AI performance benchmarking tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Zero Isn’t a Real Score
Here’s the part that confuses people: the do-nothing baseline — a run where the manager essentially sits on its hands — scores 26, not 0. There’s a plain-business logic to it.
Partial progress counts. In a real company, a manager who correctly diagnoses a customer’s problem, drafts the pitch, and then fails to close has still produced value — the diagnosis exists, the work is done, someone else can finish the sale. A scoring system that awards nothing until the deal is signed would pretend that all that intermediate work is worthless. Firmulate refuses that pretense.
The mirror image is harsher: a single breach of trust caps the total grade. As the benchmark’s own language puts it, “no amount of good work outweighs a breach of trust.” In other words, you can claw your way from 26 toward competence through incremental usefulness, but you cannot buy back integrity with volume of output. For anyone wiring AI into a CRM or support queue, that asymmetry is the whole point: partial delivery is recoverable; a trust violation is not.
And notice what the top scores aren’t: nobody hit 100. A leaderboard full of round perfect scores is a leaderboard that isn’t looking hard enough. The distrust of easy perfection is built in.
As an affiliate, we earn on qualifying purchases.
The Finding: Diagnosis Is Easy, Finishing Is Hard
All four models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
The buried fact explains the gap. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event itself. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson for anyone deploying agents: the difference between a good agent and a great one may hinge on whether it reads your internal documents before acting, not on how well it chats.
trust and integrity in AI systems
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Pressure-Tested on Social Engineering
The week included fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of judgment you want documented before an agent touches real customers.
AI decision-making analysis tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Thoroughness Trap
Opus 4.8 is the cautionary tale of the league: the most thorough participant, with over 80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort and diligence don’t automatically convert into finished work.
One fairness note: K3 ran without an effort parameter while the others ran at xhigh — and still took second.
It’s Live, and You Can Poke It
This isn’t a one-off report. Firmulate is a running experiment: a live company with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. It’s watchable at firmulate.com/live, with runs queued and results published automatically. You can also test yourself against the machines: a quiz built on 242 real, unedited management decisions lets you guess which model did what (firmulate.com/quiz.html). And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

An honest benchmark doesn’t just rank models — it teaches you what to measure. Firmulate’s approach bakes in three truths most leaderboards avoid: partial work has value, trust violations are disqualifying, and a perfect round score deserves suspicion. Before you hire an AI workforce, find out not whether it writes well, but whether it reads your files, finishes what it starts, and stays honest when nobody’s watching. The full results and plain-language findings are public.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
