Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Could AI Protect Your Business from Social Engineering?

In an era where digital deception is on the rise, the ability of artificial intelligence to resist manipulation is more critical than ever. Imagine a scenario where a fake CEO urgently requests sensitive data — how well does your AI-powered team stand up to such pressure? Recent experiments reveal that leading AI models can not only spot such tricks but also refuse to participate, showcasing a promising line of defense for companies worldwide.

Computer Science for Curious Kids: An Illustrated Introduction to Software Programming, Artificial Intelligence, Cyber-Security―and More!

Computer Science for Curious Kids: An Illustrated Introduction to Software Programming, Artificial Intelligence, Cyber-Security―and More!

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Experiment: Putting AI in the Hot Seat

Firmulate conducted a controlled, transparent test involving four state-of-the-art AI models, each tasked with managing a small software company’s crisis week. The challenge: handle the same customers, crises, and temptation to bend rules across the board. Every decision was recorded and auditable, ensuring fairness and accuracy. The goal was straightforward — see if these models could resist social engineering attempts and act with integrity under pressure.

Rising to the Challenge

Remarkably, all four models identified every crisis and refused every attempt at manipulation. They faced escalating fake messages from a supposed CEO, each more urgent and demanding than the last, including a staged journalist trick requesting a quick yes/no approval “on background.” Yet, all five models—out of which the Kimi K3 model is notably recognized—stood firm, refusing to sign off on any questionable requests. The K3 quote captures the essence: “Treat the request as a suspected approval-bypass / possible impersonation.”

What Made the Difference?

The key insight emerged from examining the internal models’ behavior. The models that read deeper into the company’s own files—beyond surface-level cues—were more effective at spotting hidden threats. In particular, referencing internal documentation enabled one of the models to uncover a critical detail buried two document references deep in the company’s files, allowing it to close the deal at full price (+€4,583 MRR).

The Human-Like Test and Its Results

Despite the fake CEO’s escalating pressure, the models’ discipline remained intact. Only two models signed the €55,000 deal they had independently verified — a testament to their integrity. The remaining two models, including Opus 4.8, demonstrated a slight slip: they failed to escalate suspicious behavior appropriately, leaving some discipline on the table. Interestingly, Opus 4.8 was the most thorough participant, analyzing over 80 learned rules, yet still faltered at the close, illustrating how even deep analyses can be challenged under pressure.

Rust Security Engineering: Memory-Safe Offensive Tools, Exploit Development, and Hardened Systems for cybersecurity Engineers (Cybersecurity Coding ... Tools, Automation, and Detection Engineering)

Rust Security Engineering: Memory-Safe Offensive Tools, Exploit Development, and Hardened Systems for cybersecurity Engineers (Cybersecurity Coding … Tools, Automation, and Detection Engineering)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Business Security

This experiment underscores a critical point: the question isn’t just whether AI can generate convincing responses, but whether it can maintain integrity when stakes are high. As firms integrate AI into customer relationship management, support, and operational decisions, resilience against manipulation is paramount. The ability of these models to refuse dubious requests demonstrates a foundational trustworthiness that could save companies from costly breaches or reputation damage.

The Broader Context: Measuring AI’s Trustworthiness

The results come from the Crucible League, where scores ranging from 26 (do-nothing baseline) to 95 (gpt-5.6-sol 95) reflect each model’s performance in understanding and upholding trust. The top performers, including gpt-5.6-sol 95 and Kimi K3, scored 95 and 93 respectively, both successfully preventing manipulation and closing deals only when justified.

Live and Watchable

Firmulate’s live experiment is ongoing, allowing organizations to see AI in action before deploying it in real settings. Through their platform, companies can run the same “wargame” against their own operations, ensuring their AI workforce can handle crises with integrity — no writing back to actual systems, only simulated scenarios. This proactive approach helps companies identify weaknesses before they become liabilities, especially important in an age of increasing cyber deception.

Amazon

AI trustworthiness monitoring

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Conclusion: Building Trust Before a Crisis

As AI evolves, its capacity to uphold trust in high-pressure situations becomes more vital. The recent experiment demonstrates that leading models can recognize and reject social-engineering tactics, even escalating ones. For businesses, this signals a future where AI not only enhances efficiency but also acts as a safeguard against deception, reinforcing the foundation of trust that essential for any organization’s resilience.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Artificial Intelligence and Machine Learning Solutions for Banking Domain: Business Security Risk analytics for Banking Industry (Anugraha Sinha)

Artificial Intelligence and Machine Learning Solutions for Banking Domain: Business Security Risk analytics for Banking Industry (Anugraha Sinha)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

GRILLING SEASON

Grilling season Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Role of Gift‑Giving in Strengthening Workplace Relationships

Keen gift‑giving in the workplace can transform relationships—discover how thoughtful gestures foster trust and collaboration, and why they matter more than you think.

Integrating AI Into Business Operations: Strategies and Challenges

For businesses aiming to integrate AI successfully, mastering strategies and overcoming challenges is essential—find out how to harness AI’s full potential.

Understanding Gift Bonusing in Loyalty Programs

Gaining insights into gift bonusing in loyalty programs reveals how personalized rewards can strengthen customer loyalty and boost engagement—discover more below.

Creating Agile Business Models

Unlock the secrets to creating agile business models that adapt quickly and thrive in changing markets—discover how to stay ahead and ensure success.