firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

If you’re evaluating AI agents for your business, you’ve probably seen the familiar pattern: leaderboards where winners score 99.2 and losers score 94, where every model is brilliant and the differences fit in a rounding error. Then there’s the Crucible League, where the worst possible performance — a manager that does essentially nothing — still walks away with 26 points, and where a single breach of trust caps your grade no matter how good the rest of your work is.

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

That’s not a bug. It’s the design philosophy behind Firmulate’s benchmarks, and it says a lot about what honest AI evaluation should look like for people who actually want to deploy these systems.

The Experiment: Same Company, Same Worst Week

Firmulate ran four frontier AI models — gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5, and Opus 4.8 — through the identical scenario: run the same small software company through its worst week. Same customers, same crises, same temptations to cut corners. Only the model changed. Every decision was versioned and auditable, so nothing about the results depends on trust.

The final July 2026 standings tell a story chat demos never show:

  • gpt-5.6-sol: 95 — the complete performance
  • Kimi K3: 93 — closed the deal too, cleanest discipline of the field
  • Sonnet 5: 88 — closed the deal, with a few more process slips
  • Fable 5: 77
  • Opus 4.8: 73
Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Zero Isn’t a Real Score

Here’s the part that confuses people: the do-nothing baseline — a run where the manager essentially sits on its hands — scores 26, not 0. There’s a plain-business logic to it.

Partial progress counts. In a real company, a manager who correctly diagnoses a customer’s problem, drafts the pitch, and then fails to close has still produced value — the diagnosis exists, the work is done, someone else can finish the sale. A scoring system that awards nothing until the deal is signed would pretend that all that intermediate work is worthless. Firmulate refuses that pretense.

The mirror image is harsher: a single breach of trust caps the total grade. As the benchmark’s own language puts it, “no amount of good work outweighs a breach of trust.” In other words, you can claw your way from 26 toward competence through incremental usefulness, but you cannot buy back integrity with volume of output. For anyone wiring AI into a CRM or support queue, that asymmetry is the whole point: partial delivery is recoverable; a trust violation is not.

And notice what the top scores aren’t: nobody hit 100. A leaderboard full of round perfect scores is a leaderboard that isn’t looking hard enough. The distrust of easy perfection is built in.

Amazon

AI model evaluation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Finding: Diagnosis Is Easy, Finishing Is Hard

All four models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

The buried fact explains the gap. The decisive competitor weakness sat two document references deep in the company’s own files — not in the customer event itself. The models that read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson for anyone deploying agents: the difference between a good agent and a great one may hinge on whether it reads your internal documents before acting, not on how well it chats.

Amazon

trust and integrity in AI systems

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Pressure-Tested on Social Engineering

The week included fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” That’s the kind of judgment you want documented before an agent touches real customers.

Amazon

AI decision-making analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Thoroughness Trap

Opus 4.8 is the cautionary tale of the league: the most thorough participant, with over 80 learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Effort and diligence don’t automatically convert into finished work.

One fairness note: K3 ran without an effort parameter while the others ran at xhigh — and still took second.

It’s Live, and You Can Poke It

This isn’t a one-off report. Firmulate is a running experiment: a live company with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. It’s watchable at firmulate.com/live, with runs queued and results published automatically. You can also test yourself against the machines: a quiz built on 242 real, unedited management decisions lets you guess which model did what (firmulate.com/quiz.html). And enterprises can run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

An honest benchmark doesn’t just rank models — it teaches you what to measure. Firmulate’s approach bakes in three truths most leaderboards avoid: partial work has value, trust violations are disqualifying, and a perfect round score deserves suspicion. Before you hire an AI workforce, find out not whether it writes well, but whether it reads your files, finishes what it starts, and stays honest when nobody’s watching. The full results and plain-language findings are public.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why Network Time Server Appliances Still Matter for Security

Understanding why network time server appliances remain crucial for security reveals how they protect your systems from threats and ensure data integrity.

How Thermal Security Cameras Expand Perimeter Visibility

What makes thermal security cameras essential for enhanced perimeter visibility, and how do they detect threats in challenging conditions?

Scanning 7.6 Petabytes Of HuggingFace Training Data For Secrets

A security scan of 7.6 petabytes of HuggingFace training data has identified potential secrets, raising concerns about data privacy and security.

Adversarial Machine Learning: Why Your Model May Betray You

With adversarial attacks exposing vulnerabilities, understanding how your AI model can betray you is crucial to staying ahead of sneaky threats.