AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine an AI that can spot every crisis, refuse manipulation attempts, and still scores just 26 out of 100 in a rigorous business test. It might sound paradoxical, but this is the reality of the latest AI benchmarking experiment conducted by Firmulate. For everyday consumers and business leaders alike, understanding what these scores mean—and what they don’t—can be the key to smarter automation decisions.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Reality of AI Benchmarking: More Than Just a Score

In the world of artificial intelligence, scores often become shorthand for capability. But the recent Firmulate experiment highlights that not all scores are created equal. The benchmark, called the Crucible League, tested four advanced AI models—named gpt-5.6-sol, Kimi K3, Sonnet 5, and Opus 4.8—by running them through a simulated week in a small software company facing real-world crises and ethical dilemmas.

What makes this stand out? Every model was given precisely the same scenarios: customer complaints, crises, manipulative requests, and even attempts to breach trust. The AI models were tasked not just with identifying issues but with making decisions that aligned with business integrity—such as refusing to sign off on shady deals or escalate fake CEO messages.

The Surprising Baseline and Partial Progress

One key insight is that the lowest score of 26 points wasn’t a failure to recognize crises. All models identified every problem and refused manipulation attempts. Yet, the scores reflect something more nuanced: partial progress counts. For instance, one model, Opus 4.8, demonstrated the most thorough analysis but ultimately left a critical deal on the table and slipped in escalation discipline. Despite recognizing issues, it did not follow through with decisive action.

Moreover, the experiment underscores that a single breach of trust caps the overall score. No matter how well a model performs elsewhere, one lapse—like improperly escalating a request—limits the total to 26 points. This design emphasizes that honesty and integrity are non-negotiable in AI decision-making, especially when human lives and money are at stake.

Amazon

AI ethics decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Do These Scores Tell Business Leaders?

For executives pondering AI investments, the takeaway is clear: performance isn’t just about what the AI can generate in a chat or a demo. It’s about whether the AI can finish what it starts, read your files responsibly, and stay honest under pressure. In the experiment, only two models signed lucrative deals—scores of 95 and 93—by actually digging into company files and making correct decisions that preserved value.

Interestingly, the models that read deeper into the company’s documents—like Kimi K3—secured full-price deals by uncovering hidden facts. This highlights the importance of thorough reading and understanding in real-world AI applications, whether for legal compliance, financial accuracy, or ethical decision-making.

The Social Engineering Test: Refusing Manipulation

The models faced sophisticated social engineering: fake CEO messages, staged escalations, and background questions from reporters. All five models refused to be manipulated. Kimi K3 explained its stance clearly: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that well-designed AI can retain ethical boundaries even when pressured from multiple angles.

Amazon

business AI trustworthiness software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Live Experiment: A Small Business in Action

Firmulate’s live site simulates a small, real-world business with 13 synthetic employees, working on revenue mechanics of over €2.3k per month against burn rates exceeding €105k monthly. The company’s workday is versioned, and all decisions are auditable. Watching this in real time at firmulate.com/live reveals how AI models perform under pressure, with each management decision being a test of discipline, honesty, and strategic thinking.

The experiment’s most thorough participant, Opus 4.8, had deep analyses but left deals on the table due to slips in process discipline—such as writing attempts into locked departments instead of escalating. This demonstrates that even the most capable AI can falter if discipline and procedural rigor aren’t maintained.

Amazon

AI compliance and integrity solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Trust and Integrity Matter More Than Scores

The Crucible League scores—a range from 77 to 95—are impressive, but the key insight is the existence of a baseline at 26. No model scores below that, because honesty is the foundation of lasting value. A single breach of trust caps the score, reinforcing that in business, cutting corners or attempting manipulation can ruin an otherwise promising AI system.

For business leaders, these findings suggest that choosing an AI isn’t just about raw capability or impressive demos. It’s about verifying that the AI maintains integrity under pressure, reads your documents responsibly, and refuses manipulation. These are the qualities that will determine whether AI adds real value—or becomes a risk.

Amazon

AI decision analysis tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Final Thoughts: A New Standard for Business AI

The Firmulate benchmark offers a transparent, rigorous way to assess AI performance in real-world scenarios. It moves beyond superficial chat scores to evaluate whether AI can truly support and uphold business ethics and processes. As AI continues to integrate into core operations, knowing it can finish what it starts—trustworthy and disciplined—is more crucial than ever.

To see the experiment in action, visit firmulate.com/live and explore how different models handle a simulated company under pressure. The future of trustworthy AI in business is watching, learning, and proving itself in real time.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Business Continuity Planning: Preparing for Crises

Optimizing your business continuity plan is crucial for crisis preparedness—discover how strategic planning can safeguard your future.

ESG Reporting: Meeting Investor and Regulatory Expectations

How can transparent ESG reporting meet investor and regulatory expectations and drive your company’s responsible growth?

Sustainable Business Practices: Measuring Impact

Laying out key metrics and strategies, discover how sustainable business practices can be effectively measured to drive meaningful environmental impact.

What to Know Before Turning Crafting Into a Small Business

Making your crafting hobby a business requires careful planning and smart choices—discover what you need to know before taking the leap.