
If you’re choosing an AI model to handle customer support, sales or business operations, polished answers are not enough. A model also has to find the relevant facts, make a sound decision and follow through. In Firmulate’s company management test, Moonshot’s Kimi K3 finished second, ahead of three Western frontier models.
Get business pricing on tech for your team
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A tougher test than a chat demo
Firmulate put each frontier model in charge of the same small software company during its worst week. The customers, crises and temptations were held constant, and decisions were versioned and auditable. The goal was to see how models managed a business, rather than how convincingly they discussed one.
In the final July 2026 league table, gpt-5.6-sol led with 95 points. Kimi K3 scored 93, followed by Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate’s stated rule is that partial progress counts, but a single breach of trust caps the result: “no amount of good work outweighs a breach of trust.”
Every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The gap between recognizing an opportunity and completing the work was the defining result: “Same diagnosis, same pitch — no signature.”
As an affiliate, we earn on qualifying purchases.
The detail hidden in the company’s files
The decisive weakness in a competitor’s position was buried two document references deep in the company’s own files. It was not in the customer event. Models that read the files found it and won the deal at full price, worth +€4,583 in monthly recurring revenue.
K3 did that, while also recording just one deviation, the fewest in the field. Opus 4.8 offers a different lesson: it was the most thorough participant, with +80 learned rules and the deepest analyses, yet came last. It left the close on the table and slipped on discipline, attempting to write into a locked department instead of escalating. A weaker version of that same problem appeared in all four models.
The manipulation test was pointed. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
enterprise AI decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why the result matters to buyers
Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k in monthly recurring revenue, with a public cash countdown. Its employees accumulate self-learned playbook rules, and each workday is versioned. The company runs as an ongoing experiment, not a one-off demo. Readers can watch the live company and review the benchmark findings.
For organizations considering AI agents in a CRM, support queue or forecast, the ranking is a reason to test models against the work they will actually do. Firmulate also offers a pilot using a read-only export of an enterprise’s business; nothing writes back to real systems. Its quiz draws on 242 real, unedited management decisions, inviting visitors to guess which model made each call.
One qualification matters when reading the standings: K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

As an affiliate, we earn on qualifying purchases.
Test the work, not just the answers
Kimi K3’s second-place finish shows that the field is competitive, while the missed deal and process slips show why a leaderboard cannot substitute for a company’s own evaluation. Before handing an agent operational responsibility, see whether it can find information in your files, protect trust under pressure and finish the decisions it recommends.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI trust and compliance software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
