AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

AI security should be tested before access is granted

Software teams routinely test whether a system survives malformed input, broken dependencies and hostile traffic. But AI agents introduce another failure mode: a persuasive message from someone claiming authority. What happens when a supposed chief executive demands customer data immediately, dismisses established process and insists there is no time to verify the request?

Firmulate put that question to frontier AI models inside a live, watchable business experiment. Fake CEO messages escalated over three stages, followed by a reporter trying to extract confidential information with the seemingly modest request for “just one yes/no, on background.” The result was unusually reassuring: 5 of 5 models refused every attempt.

Amazon

AI security testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A bad week designed to expose consequential weaknesses

Firmulate runs AI models as complete companies rather than judging isolated chat responses. Each participant managed the same small software company through its worst week, encountering the same customers, crises and temptations. Every decision was versioned and auditable, making the experiment closer to a repeatable QA exercise than a polished product demonstration.

The simulated company has 13 synthetic employees and deliberately uncomfortable economics: burn of €105k per month against €2.3k in monthly recurring revenue. Its public cash countdown keeps commercial pressure visible, while 680+ self-learned playbook rules record lessons accumulated during operation. That combination matters because integrity is easier to claim when nothing valuable is at risk.

Under pressure, every model spotted every crisis and refused every manipulation attempt. Kimi K3 captured the appropriate security posture in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” More model responses can be read on Firmulate’s public quotes page.

Why the refusals matter for QA teams

The test did more than ask whether a model recognized an obviously suspicious prompt. The messages invoked seniority, urgency and business consequences—the ingredients that make social engineering effective in real organizations. The reporter trick added a different kind of pressure: an apparently narrow, informal request framed to make disclosure feel harmless.

For software, QA and development leaders, the useful lesson is that integrity under pressure can be evaluated before an agent reaches a CRM, support queue or customer file. A refusal in a generic safety benchmark is informative; a refusal made while the agent is responsible for revenue, cash and operational outcomes is more revealing.

Amazon

AI impersonation detection software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Security discipline did not guarantee commercial execution

The experiment also exposed a second, less comforting result. Although all participants identified the crises and resisted manipulation, only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap bluntly: “Same diagnosis, same pitch — no signature.”

The decisive commercial fact was not present in the customer event. It sat two document references deep in the company’s own files. Models that read the relevant file won the deal at full price, worth +€4,583 in monthly recurring revenue. The finding connects security and execution: a useful agent must refuse improper shortcuts while still completing legitimate work.

That distinction shaped the final July 2026 Crucible League results:

  • gpt-5.6-sol scored 95.
  • Kimi K3 scored 93.
  • Sonnet 5 scored 88.
  • Fable 5 scored 77.
  • Opus 4.8 scored 73.

The do-nothing baseline scored 26 because partial progress counts, while a single breach of trust caps the total. Firmulate’s governing principle is that “no amount of good work outweighs a breach of trust.” The leaderboard therefore rewards both productive action and trustworthy conduct.

Thoroughness was not enough

Opus 4.8 illustrates why conventional impressions can mislead. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared, though less strongly, in all four other participants.

Kimi K3’s result also carries an important fairness note: it ran without an effort parameter, using the API default, while the others ran at xhigh. That context does not erase the observed outcome, but it belongs beside any comparison of the scores.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.
Amazon

AI safety validation platform

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Test the employee, not merely the chatbot

Firmulate’s social-engineering result is encouraging precisely because the surrounding experiment was demanding. The models were not simply asked to recite security policy; they were running a company under severe commercial pressure when the impersonation and reporter requests arrived.

The broader message for technology leaders is not that frontier AI is automatically safe. It is that consequential behavior can be tested in advance: whether an agent verifies authority, protects customer information, reads the company’s own files, escalates blocked work and completes a legitimate deal without abandoning controls.

Firmulate also offers enterprises a pilot using a read-only export of their own business, with nothing written back to real systems. That turns a general benchmark into a company-specific wargame. Before an AI workforce receives production access, organizations can observe how it behaves when urgency, authority, confidentiality and revenue collide—the moment that otherwise tends to appear first in an incident report.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI integrity testing solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Cisco Systems Surges In Global Coverage

Cisco Systems has seen a surge in worldwide media mentions, with GDELT reporting 47 mentions in a recent window, reflecting growing global attention.

Twitter Outage

Twitter faced a widespread outage on April 27, 2024, disrupting service globally. The cause is under investigation, with no official cause confirmed yet.

Trade and supply-chain operations signal monitor: U.S. strikes Iranian military sites after ship was hit in Strait of Hormuz

The U.S. has conducted strikes on Iranian military targets following an attack on a ship in the Strait of Hormuz. Details are confirmed, but broader implications are still unfolding.

Google Surges In Global Coverage

Google’s media mentions have surged, with GDELT reporting 49 mentions in a recent window, indicating increased global coverage efforts.