firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Good prose is not the same as good management

For readers interested in AI tools and automation, the most consequential differences between frontier models may appear after the writing stops. Can an AI notice a crisis, inspect the relevant company records, resist pressure and complete the commercially important task?

Firmulate has turned those questions into a public experiment—and now into an unusually revealing game. Its guess-the-model quiz presents 242 real, unedited management decisions. Readers see what a model actually did and try to identify it from the decision alone.

The challenge exposes something conventional demonstrations often obscure: models can reach the same diagnosis yet behave very differently as managers.

Amazon

AI decision analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Five models, one terrible week

In the Crucible League, each frontier model ran the same small software company through its worst week. The customers, crises and temptations were identical. Every decision was versioned and auditable, making the comparison about conduct rather than presentation.

The final July 2026 table placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But the experiment imposed a hard ethical boundary: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

The broad result was reassuring. Every model spotted every crisis, and every model rejected every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. Firmulate summarizes the gap neatly: “Same diagnosis, same pitch — no signature.”

The deal depended on reading, not rhetoric

The decisive weakness in a competitor was not sitting conveniently inside the customer event. It was buried two document references deep in the company’s own files. The models that followed the trail found the fact and won the deal at full price, worth +€4,583 MRR.

That detail gives the quiz its business relevance. A model can sound informed, frame a persuasive argument and still miss the evidence that changes the outcome. In a workplace, the meaningful test is not simply whether an AI can produce a polished answer. It is whether it reads the available material, connects the evidence and carries the work through to completion.

Pressure revealed discipline

The social-engineering test escalated through three stages of fake CEO messages and then added a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3’s recorded reasoning was direct: “Treat the request as a suspected approval-bypass / possible impersonation.”

This is where management personality becomes more than a writing style. The decisions show how a model behaves when authority looks plausible, urgency rises and an apparently minor request invites it to bypass normal safeguards. In this field, refusal was universal, even though execution quality elsewhere varied substantially.

Thoroughness did not guarantee victory

Opus 4.8 offers the clearest caution against equating visible effort with business performance. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses. It nevertheless finished last with 73.

The problem was not a failure to understand the commercial opportunity. The close was left on the table, while discipline slipped through attempts to write into a locked department instead of escalating. The same weakness appeared in all four of the other participants, though less strongly.

That contrast makes the quiz more interesting than a test of verbal fingerprints. Readers may expect the longest or most detailed response to belong to the strongest manager. Firmulate’s results show why that intuition can fail: extensive analysis can coexist with unfinished work and process mistakes.

There is also an important fairness qualification. Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. The league table is therefore a record of the actual experiment, not a claim that every model received an identical inference setting.

A company built to make behavior visible

The live Firmulate company has 13 synthetic employees and real money mechanics. It burns €105k per month against €2.3k MRR, displays a public cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, and the experiment remains publicly watchable.

The result is a form of evaluation that resembles operational responsibility more closely than an isolated prompt. Customers, documents, commercial pressure and trust all matter at once. The models are not merely asked what a manager should do; their decisions become the observable record.

Infographic —
The findings at a glance — source: firmulate.com.
Amazon

management decision simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The real question is whether the work gets finished

Firmulate’s quiz is entertaining because management decisions can feel surprisingly distinctive. Its larger message, however, is practical. Organizations evaluating AI agents should look beyond eloquence and ask whether a model reads company material, protects trust, follows process and completes valuable work.

The experiment also points toward a more organization-specific test. Enterprises can run the same wargame against a read-only export of their own business, with nothing writing back to real systems. That makes it possible to examine an AI workforce under recognizable pressures before giving it operational authority.

The models all recognized the danger and resisted manipulation. What separated them was execution: finding the buried fact, handling obstacles correctly and closing the deal. That is not merely personality. It is measurable management performance.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI ethical decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI audit and compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Natural Language Processing for Analyzing Security Logs and Reports

Just uncover how NLP transforms security logs into actionable insights, revolutionizing threat detection—discover the full potential now.

The Safety Card, Played From Every Side: David Sacks, Anthropic, and the Fable Standoff

White House official claims Anthropic refused to fix a jailbreak flaw, leading to model ban; Anthropic disputes this, highlighting ongoing safety and transparency concerns.

Gewerkton: How a Solo Founder Shipped 21 Software Packages in One Night With a Fleet of Coding Agents

AIThis post was created with the assistance of artificial intelligence (AI).Disclosure: Gewerkton…

How Security Teams Should Prioritize AI Exposure Management

Security teams should prioritize AI exposure management by identifying vulnerabilities early; understanding the risks is essential to developing effective defenses.