firmulate.com/live.html — live view
Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

Imagine watching a company built entirely around AI models, operating in real time, making decisions under pressure, and fighting to survive — all in full view. This is the extreme experiment of Firmulate, where a simulated business with no human employees is tested daily against crises, temptations, and tough negotiations.

At a glance, it’s a typical tech experiment: a small software company run by artificial intelligence models, with zero human staff. But look closer, and you’ll find a unique, unfiltered window into how AI can manage a business — or struggle to do so. This is the live experiment of Firmulate, a platform that puts AI models through the grind of running a real company, complete with real money mechanics, public cash countdowns, and a daily versioned decision log.

Every workday, the system operates with 13 synthetic employees — each modeled to perform specific roles — and faces the same crises and temptations as any human-run business. The company burns €105,000 each month but earns just €2,300 in monthly recurring revenue, creating a stark financial reality that underscores the urgency and stakes of the project.

This setup is not just for show. It’s a carefully constructed battleground where AI models are tested on their judgment, discipline, and honesty. The core question: can these models not only spot every crisis but also resist manipulation attempts, especially in high-pressure social engineering scenarios like fake CEO messages or reporter tricks? And more critically, can they close deals and generate revenue under these conditions?

Recent results from the experiment reveal both promise and limitations. Four frontier models, including GPT-5.6-SOL, Kimi K3, Sonnet 5, and Fable 5, were each tasked with running the same tough week. All identified every crisis and refused every manipulation attempt, even as social engineering attempts escalated over multiple stages. That’s promising: the models showed resilience against deception. But where things got interesting was in closing deals.

Only two models, GPT-5.6-SOL and Kimi K3, managed to sign the €55,000 deal their analyses had recommended — but only after reading deep into the company’s own files to uncover a hidden, decisive fact buried two document references deep. That crucial insight was the key to winning the full price, adding over €4,500 in monthly recurring revenue. The other two models failed to follow through, leaving potential revenue unclaimed, despite accurate diagnoses.

This discrepancy highlights a critical blind spot for current AI models: reading comprehension extends beyond surface interactions. The models that examined internal documents thoroughly had a tangible advantage, emphasizing the importance of deep contextual understanding in real-world business tasks.

Beyond decision accuracy, the experiment also tested social engineering defenses. Fake CEO messages escalating over three stages and a reporter trick — asking for a simple yes/no on background — were both refused by all models. Kimi K3’s reasoning was clear: treat such requests as suspects of impersonation or approval bypass. This indicates growing robustness against manipulation tactics that often fool human decision-makers.

The live system at firmulate.com offers a real-time view of this ongoing effort. It features a continuously evolving setup where every decision is versioned, and the entire process is transparent to viewers. As of now, the experiment operates with a total burn rate of €105,000 per month against a tiny monthly revenue stream, making the company’s survival a race against time. The site also displays a leaderboard of models, with GPT-5.6-SOL leading at 95 points, followed closely by Kimi K3 at 93, and others trailing behind.

While some models excelled at rule discipline and diagnosis, others struggled with closing deals — most notably Opus 4.8, which demonstrated the most thorough analysis but failed to seize opportunities, leaving potential revenue on the table. The experiment underscores an essential truth for AI in management: understanding isn’t enough; execution and discipline matter just as much.

For enterprise leaders, the takeaway is clear. When deploying AI agents for tasks like CRM, customer support, or forecasting, their ability to see through crises, maintain honesty, and follow through on deals is crucial. It’s not just about how well an AI writes or chats, but whether it can reliably complete work, uphold integrity under pressure, and ultimately drive value.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.

Firmulate’s live experiment shows AI models can detect crises and resist manipulation, but closing deals and executing strategies remains a challenge — vital skills for AI business tools.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


AI Builders: Making The Decisions That Turn AI Code Into Real Software

AI Builders: Making The Decisions That Turn AI Code Into Real Software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

AI for Public Relations: A How-To Guide for Implementation and Management

AI for Public Relations: A How-To Guide for Implementation and Management

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Never Be Closing: Surviving and Thriving in the AI Apocalypse with Culture, Speed, and Soul.

Never Be Closing: Surviving and Thriving in the AI Apocalypse with Culture, Speed, and Soul.

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)

Practical Claude Handbook for Attorneys: Master Case Analysis, Contract Review, Research Automation, Client Communication, and Document Drafting (Claude AI Guide for Beginners)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Strengthening Password Security With AI: Behavioral Biometrics and More

Keen on enhancing your password security? Discover how AI-driven behavioral biometrics can keep your accounts safer than ever.

Predictive Threat Intelligence: Can AI Really See Tomorrow’s Attack?

What if AI could forecast cyber threats before they happen, but how reliable is this glimpse into tomorrow’s attacks? Keep reading to find out.

How AI Stopped the Biggest Data Breach in History – You Won't Believe How

Preventing catastrophic consequences, AI's lightning-fast response to the massive breach left experts stunned, but how did it achieve the impossible?

Ai-Powered Firewall Learns From Every Attack – Now Unhackable?

Sophisticated AI-powered firewalls autonomously evolve defenses, but can they truly guarantee unhackable security in the face of relentless cyber threats?