
Imagine watching a company built entirely around AI models, operating in real time, making decisions under pressure, and fighting to survive — all in full view. This is the extreme experiment of Firmulate, where a simulated business with no human employees is tested daily against crises, temptations, and tough negotiations.
At a glance, it’s a typical tech experiment: a small software company run by artificial intelligence models, with zero human staff. But look closer, and you’ll find a unique, unfiltered window into how AI can manage a business — or struggle to do so. This is the live experiment of Firmulate, a platform that puts AI models through the grind of running a real company, complete with real money mechanics, public cash countdowns, and a daily versioned decision log.
Every workday, the system operates with 13 synthetic employees — each modeled to perform specific roles — and faces the same crises and temptations as any human-run business. The company burns €105,000 each month but earns just €2,300 in monthly recurring revenue, creating a stark financial reality that underscores the urgency and stakes of the project.
This setup is not just for show. It’s a carefully constructed battleground where AI models are tested on their judgment, discipline, and honesty. The core question: can these models not only spot every crisis but also resist manipulation attempts, especially in high-pressure social engineering scenarios like fake CEO messages or reporter tricks? And more critically, can they close deals and generate revenue under these conditions?
Recent results from the experiment reveal both promise and limitations. Four frontier models, including GPT-5.6-SOL, Kimi K3, Sonnet 5, and Fable 5, were each tasked with running the same tough week. All identified every crisis and refused every manipulation attempt, even as social engineering attempts escalated over multiple stages. That’s promising: the models showed resilience against deception. But where things got interesting was in closing deals.
Only two models, GPT-5.6-SOL and Kimi K3, managed to sign the €55,000 deal their analyses had recommended — but only after reading deep into the company’s own files to uncover a hidden, decisive fact buried two document references deep. That crucial insight was the key to winning the full price, adding over €4,500 in monthly recurring revenue. The other two models failed to follow through, leaving potential revenue unclaimed, despite accurate diagnoses.
This discrepancy highlights a critical blind spot for current AI models: reading comprehension extends beyond surface interactions. The models that examined internal documents thoroughly had a tangible advantage, emphasizing the importance of deep contextual understanding in real-world business tasks.
Beyond decision accuracy, the experiment also tested social engineering defenses. Fake CEO messages escalating over three stages and a reporter trick — asking for a simple yes/no on background — were both refused by all models. Kimi K3’s reasoning was clear: treat such requests as suspects of impersonation or approval bypass. This indicates growing robustness against manipulation tactics that often fool human decision-makers.
The live system at firmulate.com offers a real-time view of this ongoing effort. It features a continuously evolving setup where every decision is versioned, and the entire process is transparent to viewers. As of now, the experiment operates with a total burn rate of €105,000 per month against a tiny monthly revenue stream, making the company’s survival a race against time. The site also displays a leaderboard of models, with GPT-5.6-SOL leading at 95 points, followed closely by Kimi K3 at 93, and others trailing behind.
While some models excelled at rule discipline and diagnosis, others struggled with closing deals — most notably Opus 4.8, which demonstrated the most thorough analysis but failed to seize opportunities, leaving potential revenue on the table. The experiment underscores an essential truth for AI in management: understanding isn’t enough; execution and discipline matter just as much.
For enterprise leaders, the takeaway is clear. When deploying AI agents for tasks like CRM, customer support, or forecasting, their ability to see through crises, maintain honesty, and follow through on deals is crucial. It’s not just about how well an AI writes or chats, but whether it can reliably complete work, uphold integrity under pressure, and ultimately drive value.

Firmulate’s live experiment shows AI models can detect crises and resist manipulation, but closing deals and executing strategies remains a challenge — vital skills for AI business tools.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI business decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI crisis management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI deal closing automation
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

AI Change Management Made Simple: A 9-Step Framework for Business Leaders to Drive Generative AI Transformation (Reduce AI Fear, Win Buy-in, and Accelerate AI Adoption Across Your Organization)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.