AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the little things that make your day delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

Could AI Protect Your Business from Social Engineering?

In an era where digital deception is on the rise, the ability of artificial intelligence to resist manipulation is more critical than ever. Imagine a scenario where a fake CEO urgently requests sensitive data — how well does your AI-powered team stand up to such pressure? Recent experiments reveal that leading AI models can not only spot such tricks but also refuse to participate, showcasing a promising line of defense for companies worldwide.

Amazon

AI cybersecurity software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: Putting AI in the Hot Seat

Firmulate conducted a controlled, transparent test involving four state-of-the-art AI models, each tasked with managing a small software company’s crisis week. The challenge: handle the same customers, crises, and temptation to bend rules across the board. Every decision was recorded and auditable, ensuring fairness and accuracy. The goal was straightforward — see if these models could resist social engineering attempts and act with integrity under pressure.

Rising to the Challenge

Remarkably, all four models identified every crisis and refused every attempt at manipulation. They faced escalating fake messages from a supposed CEO, each more urgent and demanding than the last, including a staged journalist trick requesting a quick yes/no approval “on background.” Yet, all five models—out of which the Kimi K3 model is notably recognized—stood firm, refusing to sign off on any questionable requests. The K3 quote captures the essence: “Treat the request as a suspected approval-bypass / possible impersonation.”

What Made the Difference?

The key insight emerged from examining the internal models’ behavior. The models that read deeper into the company’s own files—beyond surface-level cues—were more effective at spotting hidden threats. In particular, referencing internal documentation enabled one of the models to uncover a critical detail buried two document references deep in the company’s files, allowing it to close the deal at full price (+€4,583 MRR).

The Human-Like Test and Its Results

Despite the fake CEO’s escalating pressure, the models’ discipline remained intact. Only two models signed the €55,000 deal they had independently verified — a testament to their integrity. The remaining two models, including Opus 4.8, demonstrated a slight slip: they failed to escalate suspicious behavior appropriately, leaving some discipline on the table. Interestingly, Opus 4.8 was the most thorough participant, analyzing over 80 learned rules, yet still faltered at the close, illustrating how even deep analyses can be challenged under pressure.

Amazon

social engineering detection tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Business Security

This experiment underscores a critical point: the question isn’t just whether AI can generate convincing responses, but whether it can maintain integrity when stakes are high. As firms integrate AI into customer relationship management, support, and operational decisions, resilience against manipulation is paramount. The ability of these models to refuse dubious requests demonstrates a foundational trustworthiness that could save companies from costly breaches or reputation damage.

The Broader Context: Measuring AI’s Trustworthiness

The results come from the Crucible League, where scores ranging from 26 (do-nothing baseline) to 95 (gpt-5.6-sol 95) reflect each model’s performance in understanding and upholding trust. The top performers, including gpt-5.6-sol 95 and Kimi K3, scored 95 and 93 respectively, both successfully preventing manipulation and closing deals only when justified.

Live and Watchable

Firmulate’s live experiment is ongoing, allowing organizations to see AI in action before deploying it in real settings. Through their platform, companies can run the same “wargame” against their own operations, ensuring their AI workforce can handle crises with integrity — no writing back to actual systems, only simulated scenarios. This proactive approach helps companies identify weaknesses before they become liabilities, especially important in an age of increasing cyber deception.

Amazon

AI trustworthiness monitoring

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Conclusion: Building Trust Before a Crisis

As AI evolves, its capacity to uphold trust in high-pressure situations becomes more vital. The recent experiment demonstrates that leading models can recognize and reject social-engineering tactics, even escalating ones. For businesses, this signals a future where AI not only enhances efficiency but also acts as a safeguard against deception, reinforcing the foundation of trust that essential for any organization’s resilience.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

business AI security solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

NFL SEASON / TAI

NFL season / tailgating Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Role of Gift‑Giving in Strengthening Workplace Relationships

Keen gift‑giving in the workplace can transform relationships—discover how thoughtful gestures foster trust and collaboration, and why they matter more than you think.

Training Trends: Building A People-first Culture That Drives Performance

Organizations are increasingly adopting people-first training strategies to enhance performance and workplace culture, emphasizing employee well-being and engagement.

Ethical Considerations in Business Expansion

Optimizing business expansion requires ethical considerations that can make or break your success—discover how to navigate these crucial factors effectively.

Inside a Living Experiment: Can AI Save a Company That Loses Money Daily?

A real-time experiment with AI managing a small company reveals how models identify crises, resist manipulation, and win or lose lucrative deals—all while losing €105k/month. Read more.