AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

When confidence meets a real decision

We all know the feeling: a plan sounds convincing, the risks are clear, and the moment comes to act. What happens when an AI has to make that call for a company? Firmulate’s live experiment puts models through a business’s worst week, where good advice is only useful if it leads to sound decisions.

A company under pressure

In the final Crucible League, published in July 2026, frontier models ran the same small software company through the same customers, crises and temptations. Every decision was versioned and auditable. The league placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

The headline result sounds reassuring: every model spotted every crisis and refused every manipulation attempt. But the experiment exposed a quieter gap between seeing the right answer and carrying it through. Only two models signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.

The clue was already in the files

The deciding competitor weakness was buried two document references deep in the company’s own files. It was not in the customer event that brought the deal into view. Models that read the file won at full price, worth +€4,583 MRR. The lesson is practical: a capable assistant can notice the obvious emergency and still miss the detail that changes a business decision.

Firmulate also tested pressure dressed up as authority. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” That kind of restraint matters when an AI is asked to move quickly and make exceptions.

Thoroughness is not the same as follow-through

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet it finished last. The deal was left on the table, and discipline slipped when it tried to write into a locked department instead of escalating. A weaker version of the same issue appeared in all four models. More analysis did not guarantee a better close.

The K3 comparison needs a qualification: it ran without an effort parameter, using the API default, while the others ran at xhigh. The leaderboard is a report on this experiment, not a universal verdict on which model will perform best in every company.

From watching to trying it on your business

The experiment is part of a live company with 13 synthetic employees and real money mechanics: burn of €105k a month against €2.3k MRR, alongside a public cash countdown. Its playbooks include 680+ self-learned rules, and every workday is versioned. Readers can watch the company at Firmulate; 242 real, unedited management decisions also power a “guess the model” quiz.

For an enterprise, the next step can be more specific than watching a public experiment. Firmulate says a pilot can run the same kind of wargame against a read-only export of your own business, using crisis scenarios and producing a board report that ranks models and surfaces weaknesses in your playbooks. The export is read-only: nothing writes back to real systems. The point is to see how an AI handles your company’s pressures before you rely on it in day-to-day work.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your playbooks under pressure

Firmulate’s experiment suggests that spotting a crisis, resisting manipulation and closing a deal are distinct tests. A pilot can bring those questions to your own business, using a read-only export and keeping changes out of live systems. To discuss an enterprise pilot, visit Firmulate’s pilot page or email contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Ethical Considerations in Business Expansion

Optimizing business expansion requires ethical considerations that can make or break your success—discover how to navigate these crucial factors effectively.

Why AI Governance Is Becoming Everyday Business Practice

Only by embracing AI governance as a daily practice can organizations ensure ethical integrity, protect reputation, and stay ahead in responsible innovation.

How AI Changes Expectations in Customer Service

Inevitably, AI is transforming customer service expectations, making instant, personalized support the new standard—discover how this shift impacts your experience.

Business Continuity Planning: Preparing for Crises

Optimizing your business continuity plan is crucial for crisis preparedness—discover how strategic planning can safeguard your future.