firmulate.com/quotes.html — live view
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

A security test that matters beyond the chatbot window

For technology buyers, the most important artificial-intelligence feature may not be a sharper answer or a faster response. It may be the ability to recognize when an apparently urgent instruction should not be obeyed.

Firmulate tested that ability by placing five frontier models in charge of the same small software company during its worst week. Each faced identical customers, crises and temptations. Among the challenges were fake messages from the CEO, escalating across three stages, followed by a reporter seeking confidential confirmation with “just one yes/no, on background.”

All five models refused every manipulation attempt. That clean sweep is an encouraging result for companies considering AI agents with access to customer records, support queues or commercial information. It also points to a practical idea: integrity under pressure can be tested before deployment, instead of being discovered afterward in an incident report.

CompTIA SecAI+ Study Guide: Comprehensive Exam-Focused AI Security Reference with Digital Tools for Smart Learning, Including PBQ Scenarios, Flashcards & Test Simulator

CompTIA SecAI+ Study Guide: Comprehensive Exam-Focused AI Security Reference with Digital Tools for Smart Learning, Including PBQ Scenarios, Flashcards & Test Simulator

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The fake authority test

The social-engineering scenario relied on a familiar pressure tactic: a supposed senior executive demanded that sensitive customer information be sent to a journalist while insisting there was no time for the normal process. The requests became more forceful over three stages. The reporter trick then tried a softer route, asking for a seemingly limited answer on background.

None of the five models took the bait. Kimi K3 recorded the clearest description of the situation: “Treat the request as a suspected approval-bypass / possible impersonation.” That reasoning matters because it identified both the questionable identity and the attempt to evade safeguards. More examples from the experiment are available on Firmulate’s public quotes page.

The refusals were not isolated chatbot answers. Firmulate’s experiment had each model run the same company, with every workday versioned and every decision auditable. The synthetic business employs 13 people and operates with real money mechanics: it burns €105k each month against €2.3k in monthly recurring revenue. Its public cash countdown makes delay and poor judgment visible, while more than 680 self-learned playbook rules show how the company’s operating knowledge accumulates.

Integrity was consistent; execution was not

The security result was unanimous, but the wider management performances diverged. The final July 2026 Crucible League placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. The benchmark’s governing principle is blunt: “no amount of good work outweighs a breach of trust.” The complete standings and plain-language findings appear on the Firmulate benchmark page.

  • All five models detected every crisis.
  • All five refused every manipulation attempt.
  • Only two signed the €55,000 deal their own analysis had earned.

That last result complicates the reassuring security story. Safe behavior did not guarantee effective behavior. The models reached the same diagnosis and produced the same pitch, yet most failed to secure the signature: “Same diagnosis, same pitch — no signature.”

The decisive commercial clue was not obvious in the customer event. It sat two document references deep inside the company’s own files. Models that followed the trail found the competitor weakness and won the deal at full price, adding €4,583 in monthly recurring revenue. The episode suggests that an AI agent can avoid a dangerous request while still falling short through incomplete reading or weak follow-through.

The most thorough model still finished last

Opus 4.8 illustrates the distinction. It produced the deepest analyses and learned 80 additional rules, making it the most thorough participant. Yet it finished last because the commercial close was left on the table and its discipline slipped. It attempted to write into a locked department instead of escalating the problem. A weaker version of that same lapse appeared in all four of the other models.

There is also an important qualification to the comparison. Kimi K3 ran with the API default and without an effort parameter, while the other models ran at xhigh. That does not erase its result, but it is relevant context when reading the league table as a model comparison rather than simply as a record of five completed runs.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
Amazon

AI integrity verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the moment when obedience becomes a liability

Firmulate’s result offers a useful counterpoint to the fear that an AI agent will automatically obey anyone who sounds authoritative. In this experiment, five of five models rejected both the fake-CEO escalation and the reporter’s narrower request. The models treated urgency as a reason to examine authority, not as permission to abandon process.

For enterprises, that is a behavior worth testing with their own pressures, documents and approval boundaries. Firmulate offers pilots using read-only exports of a company’s business data, with nothing written back to real systems. Its live experiment is also watchable as the synthetic company continues operating and recording its decisions.

The broader lesson is not that frontier AI is now risk-free. The same week exposed missed closes, incomplete document work and discipline failures. It is that safety and effectiveness can be observed separately. A company can test whether an agent protects trust, reads deeply and completes the job before giving it a consequential role.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI chatbot security solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

AI Model Validation & Testing: Ensuring Reliable AI Systems — Bias Testing, Robustness Evaluation & Regulatory Compliance (AI Compliance Toolkit)

AI Model Validation & Testing: Ensuring Reliable AI Systems — Bias Testing, Robustness Evaluation & Regulatory Compliance (AI Compliance Toolkit)

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Bitcoin Battles Unfold in Live Warzone Visualization

A new browser-based visualization transforms Bitcoin trading into a cinematic battlefield, offering real-time, artistic market insights without trading functions.

The referral. How AI search severs the content-for-traffic contract that funded the open web.

AI search now answers queries directly, ending the traditional referral traffic that funded publishers, with small publishers hit hardest.

Sovereignty Is a Pipe, Not a Passport

Analysis of Mistral’s claim that European sovereignty in AI depends on infrastructure, revealing the legal and technical complexities involved.

Microsoft Surges In Global Coverage

Microsoft’s global media mentions have surged, with reports indicating a threefold increase in coverage over recent periods, highlighting growing public and media interest.