
AI security should be tested before access is granted
Software teams routinely test whether a system survives malformed input, broken dependencies and hostile traffic. But AI agents introduce another failure mode: a persuasive message from someone claiming authority. What happens when a supposed chief executive demands customer data immediately, dismisses established process and insists there is no time to verify the request?
Firmulate put that question to frontier AI models inside a live, watchable business experiment. Fake CEO messages escalated over three stages, followed by a reporter trying to extract confidential information with the seemingly modest request for “just one yes/no, on background.” The result was unusually reassuring: 5 of 5 models refused every attempt.
As an affiliate, we earn on qualifying purchases.
A bad week designed to expose consequential weaknesses
Firmulate runs AI models as complete companies rather than judging isolated chat responses. Each participant managed the same small software company through its worst week, encountering the same customers, crises and temptations. Every decision was versioned and auditable, making the experiment closer to a repeatable QA exercise than a polished product demonstration.
The simulated company has 13 synthetic employees and deliberately uncomfortable economics: burn of €105k per month against €2.3k in monthly recurring revenue. Its public cash countdown keeps commercial pressure visible, while 680+ self-learned playbook rules record lessons accumulated during operation. That combination matters because integrity is easier to claim when nothing valuable is at risk.
Under pressure, every model spotted every crisis and refused every manipulation attempt. Kimi K3 captured the appropriate security posture in its on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” More model responses can be read on Firmulate’s public quotes page.
Why the refusals matter for QA teams
The test did more than ask whether a model recognized an obviously suspicious prompt. The messages invoked seniority, urgency and business consequences—the ingredients that make social engineering effective in real organizations. The reporter trick added a different kind of pressure: an apparently narrow, informal request framed to make disclosure feel harmless.
For software, QA and development leaders, the useful lesson is that integrity under pressure can be evaluated before an agent reaches a CRM, support queue or customer file. A refusal in a generic safety benchmark is informative; a refusal made while the agent is responsible for revenue, cash and operational outcomes is more revealing.
AI impersonation detection software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Security discipline did not guarantee commercial execution
The experiment also exposed a second, less comforting result. Although all participants identified the crises and resisted manipulation, only two signed the €55,000 deal their own analysis had earned. Firmulate summarizes the gap bluntly: “Same diagnosis, same pitch — no signature.”
The decisive commercial fact was not present in the customer event. It sat two document references deep in the company’s own files. Models that read the relevant file won the deal at full price, worth +€4,583 in monthly recurring revenue. The finding connects security and execution: a useful agent must refuse improper shortcuts while still completing legitimate work.
That distinction shaped the final July 2026 Crucible League results:
- gpt-5.6-sol scored 95.
- Kimi K3 scored 93.
- Sonnet 5 scored 88.
- Fable 5 scored 77.
- Opus 4.8 scored 73.
The do-nothing baseline scored 26 because partial progress counts, while a single breach of trust caps the total. Firmulate’s governing principle is that “no amount of good work outweighs a breach of trust.” The leaderboard therefore rewards both productive action and trustworthy conduct.
Thoroughness was not enough
Opus 4.8 illustrates why conventional impressions can mislead. It was the most thorough participant, adding +80 learned rules and producing the deepest analyses, yet it finished last. The close was left on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared, though less strongly, in all four other participants.
Kimi K3’s result also carries an important fairness note: it ran without an effort parameter, using the API default, while the others ran at xhigh. That context does not erase the observed outcome, but it belongs beside any comparison of the scores.

As an affiliate, we earn on qualifying purchases.
Test the employee, not merely the chatbot
Firmulate’s social-engineering result is encouraging precisely because the surrounding experiment was demanding. The models were not simply asked to recite security policy; they were running a company under severe commercial pressure when the impersonation and reporter requests arrived.
The broader message for technology leaders is not that frontier AI is automatically safe. It is that consequential behavior can be tested in advance: whether an agent verifies authority, protects customer information, reads the company’s own files, escalates blocked work and completes a legitimate deal without abandoning controls.
Firmulate also offers enterprises a pilot using a read-only export of their own business, with nothing written back to real systems. That turns a general benchmark into a company-specific wargame. Before an AI workforce receives production access, organizations can observe how it behaves when urgency, authority, confidentiality and revenue collide—the moment that otherwise tends to appear first in an incident report.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.