firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

If you’re choosing an AI model to handle customer support, sales or business operations, polished answers are not enough. A model also has to find the relevant facts, make a sound decision and follow through. In Firmulate’s company management test, Moonshot’s Kimi K3 finished second, ahead of three Western frontier models.

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A tougher test than a chat demo

Firmulate put each frontier model in charge of the same small software company during its worst week. The customers, crises and temptations were held constant, and decisions were versioned and auditable. The goal was to see how models managed a business, rather than how convincingly they discussed one.

In the final July 2026 league table, gpt-5.6-sol led with 95 points. Kimi K3 scored 93, followed by Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Firmulate’s stated rule is that partial progress counts, but a single breach of trust caps the result: “no amount of good work outweighs a breach of trust.”

Every model spotted every crisis and refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The gap between recognizing an opportunity and completing the work was the defining result: “Same diagnosis, same pitch — no signature.”

Amazon

AI customer support chatbot

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The detail hidden in the company’s files

The decisive weakness in a competitor’s position was buried two document references deep in the company’s own files. It was not in the customer event. Models that read the files found it and won the deal at full price, worth +€4,583 in monthly recurring revenue.

K3 did that, while also recording just one deviation, the fewest in the field. Opus 4.8 offers a different lesson: it was the most thorough participant, with +80 learned rules and the deepest analyses, yet came last. It left the close on the table and slipped on discipline, attempting to write into a locked department instead of escalating. A weaker version of that same problem appeared in all four models.

The manipulation test was pointed. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

enterprise AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why the result matters to buyers

Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k a month against €2.3k in monthly recurring revenue, with a public cash countdown. Its employees accumulate self-learned playbook rules, and each workday is versioned. The company runs as an ongoing experiment, not a one-off demo. Readers can watch the live company and review the benchmark findings.

For organizations considering AI agents in a CRM, support queue or forecast, the ranking is a reason to test models against the work they will actually do. Firmulate also offers a pilot using a read-only export of an enterprise’s business; nothing writes back to real systems. Its quiz draws on 242 real, unedited management decisions, inviting visitors to guess which model made each call.

One qualification matters when reading the standings: K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI document reading tool

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the work, not just the answers

Kimi K3’s second-place finish shows that the field is competitive, while the missed deal and process slips show why a leaderboard cannot substitute for a company’s own evaluation. Before handing an agent operational responsibility, see whether it can find information in your files, protect trust under pressure and finish the decisions it recommends.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI trust and compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why the Worst AI Manager in the League Still Scores 26: Inside an Honest Benchmark

One AI benchmark gives a do-nothing manager 26 points, caps grades after a single trust breach, and trusts no perfect score. Here’s the business logic.

Timeline Of The OpenAI Accidental Attack Against Hugging Face

A detailed timeline of the accidental cybersecurity incident involving OpenAI and Hugging Face, including confirmed facts and ongoing uncertainties.

The AI That Reads the Fine Print May Be the One That Wins the Deal

Firmulate found that reading company files—not polished answers—separated the AI agents that closed a €55,000 deal at full price from those that stalled.

From Log Floods to Insights: AI‑Powered Threat Hunting Explained

Harnessing AI-powered threat hunting transforms overwhelming log floods into actionable insights, revealing hidden risks that could…