firmulate.com/pilot.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

For businesses adopting AI tools and automation, the hard question is no longer whether a model can spot a crisis. It is whether an AI workforce can make the right call when a real customer, a tempting shortcut and a company’s own rules collide. Firmulate’s live experiment puts models in charge of the same small software company to find out.

Buying for a business?Offer from Amazon

Get business pricing on tech for your team

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

A business stress test, not a chat demo

In the final Crucible League, published in July 2026, frontier models ran the same company through its worst week: the same customers, crises and temptations. Every decision was versioned and auditable. The results put gpt-5.6-sol first with 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. A do-nothing baseline scored 26. The league’s integrity rule is blunt: a single breach of trust caps the total. “No amount of good work outweighs a breach of trust.”

The broad finding was reassuring, but incomplete. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The diagnosis and pitch were there; the signature was not. For organizations considering autonomous agents, that gap between understanding a task and finishing it is the point of the exercise.

The clue was buried in the company’s own files

The deal turned on a competitor weakness two document references deep in the company’s files, rather than in the customer event itself. Models that read that material won the deal at full price, worth €4,583 in monthly recurring revenue. The episode shows why evaluating an agent on a tidy prompt can miss the practical challenge: important context may be scattered across the information a business already holds.

Firmulate also tested resistance to social engineering. Fake messages from a CEO escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That is the kind of judgment a company needs to see before giving an agent access to sensitive work.

Capability includes discipline

Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, but it finished last. It left the deal unsigned and tried to write into a locked department instead of escalating. A weaker version of the same discipline problem appeared in all four models. More analysis, on its own, did not guarantee a better business outcome.

The experiment runs inside a live synthetic company with 13 employees, a public cash countdown, and real money mechanics: €105,000 in monthly burn against €2,300 in monthly recurring revenue. Its playbooks contain more than 680 self-learned rules, and every workday is versioned. Readers can follow the company at Firmulate. A quiz based on 242 real, unedited management decisions also invites readers to guess which model made each call.

There is a caveat in the ranking. Kimi K3 ran without an effort parameter, using its API default, while the other models ran at xhigh. That difference belongs alongside the scores when interpreting the results.

From watching to testing your own business

For an enterprise, the next step is to test agents against the details and pressures of its own operation. Firmulate says a pilot can use a read-only export of a company’s business data to build a digital twin and run crisis scenarios, then produce a board report with model rankings and weak points in the company’s playbooks. The pilot does not write back to real systems.

That setup makes the exercise relevant to teams deciding how AI might touch a CRM, support queue or forecast: they can examine decisions under pressure before connecting agents to live operations. A simulated company cannot establish how every deployment will behave, but it can reveal where models hesitate, miss context or fail to follow through.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Firmulate’s league suggests that spotting danger and resisting manipulation are only part of the job. Closing a deal, using buried context and escalating when blocked matter too. Enterprises can explore a pilot using a read-only export of their business; nothing writes back to real systems. To discuss a pilot, contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Be Skeptical Of OpenAI’s Rogue Hacker Agent Story

Experts urge caution in accepting OpenAI’s story of a rogue hacker agent, citing lack of verified evidence and raising concerns over misinformation.

Three Key AI Warnings That Were Almost Overlooked

A detailed analysis of three overlooked AI threats, based on recent investigations into OpenAI’s security incidents and emerging agent capabilities.

Generative AI in Phishing and Scams: Emerging Threats and Solutions

Understanding how generative AI fuels evolving phishing scams reveals why staying vigilant is more critical than ever.

The Future of AI in Cybersecurity: Trends and Predictions

Looming ahead are transformative AI trends in cybersecurity that could redefine how we detect and prevent threats—discover what’s next.