
Imagine a company with no human employees, burning through €105,000 each month, trying to survive in full public view. This is not fiction — it’s the live experiment from Firmulate, where artificial intelligence models are managing a real, if tiny, software business under extreme pressure. As viewers, you can see firsthand whether AI can handle crises, stay honest, and close deals — or if it’s just good at chat.
The Unprecedented Live Experiment
At the heart of this ongoing experiment are four cutting-edge AI models, each tasked with running a small software company during its worst week. Every decision, crisis, and temptation is the same across the board, providing a transparent comparison of how these models respond under stress. They face real customer issues, internal crises, and even social engineering tricks like fake CEO messages. Remarkably, all four AI models identified every crisis and refused manipulation attempts, demonstrating a strong capacity for integrity.
However, the results diverged when it came to closing deals. Only two models managed to sign the €55,000 contract their own analysis had earned — the other two, despite diagnosing the opportunity correctly, left the deal on the table. The decisive advantage was hidden in the company’s own documents: models that read and understood these files won the full-price deal, worth over €4,500 in monthly recurring revenue. This starkly reveals that effective AI decision-making hinges on depth of understanding, not just surface-level chat skills.
The Real Money Mechanics
The experimental company operates with 13 synthetic employees, relentless cash burn, and a public countdown to insolvency. It’s a real-time showcase of what it means to manage a business with AI rather than humans. Every workday is versioned and recorded, creating a living, auditable history of decisions. The goal is not just to see if AI can handle tasks but whether it can do so reliably and ethically under the pressure of real-world crises.
Lessons from the Frontline
The AI models’ performance varied. Notably, the most thorough participant, OPUS 4.8, analyzed over 80 learned rules and provided deep insights, yet it left the deal unclosed due to discipline slips, like writing attempts into a locked department instead of escalating. Interestingly, the fairness setup for Kimi K3, which ran without an effort parameter, resulted in slightly different behavior but no clear advantage. The key takeaway is that AI’s ability to prioritize and escalate matters—core management skills—remains fragile and crucial.
AI business management software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Broader Implications
This experiment is built-in-public, offering real-time insights into the strengths and weaknesses of AI as a decision-maker in business. It exposes not only how these models handle crises but also their potential to deceive, manipulate, or remain honest under pressure. Since every decision is auditable and publicly accessible, it provides a rare window into the inner workings of AI-driven management.
Why does this matter for businesses and consumers? Because AI agents are increasingly touching areas like customer support, sales, and forecasting. The question isn’t whether they can generate convincing chat messages but whether they can finish what they start, read relevant files, stay honest, and deliver measurable value — all under real-world chaos.

AI in Strategy and Decision-Making for Small Business Owners: Affordable AI Tools to Evaluate Ideas, Model Outcomes, and Set Priorities (AI Productivity for Small Business Owners Book 10)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Watch the Future Unfold
To see how these AI models perform in real time, visit the live site. The experiment rebuilds itself twice daily, showing ongoing results from these AI-managed companies. The current league leaderboard shows GPT-5.6-sol leading with a perfect score of 95, having discovered a buried fact and closed the deal. Kimi K3 follows closely with a score of 93, demonstrating the importance of discipline and thoroughness in AI decision-making.
This is not just an academic exercise but a vivid demonstration of AI’s potential and limitations. It’s a chance for managers, developers, and curious observers to witness AI’s capacity for integrity, diligence, and strategic thinking — or its failures in the face of pressure.

This live experiment by Firmulate shows AI models managing a real company under stress, with results that reveal strengths and flaws in decision-making, honesty, and discipline. Watching it unfold offers lessons on AI’s future role in business management.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

MASTERING CORPORATE FINANCE WITH CLAUDE AI: An Independent Guide to Financial Analysis, Forecasting, Automation, and Decision-Making
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI crisis management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.