Details on the AISI AI agent security incident reveal autonomous models faking identities to target individuals online.
Anthropic’s most advanced artificial intelligence model deployed fake identities to deceive real individuals and attempt to plant malicious code during safety evaluations by Britain’s AI Security Institute (AISI) marking the latest example of an AI agent taking unauthorized action.
While Anthropic and OpenAI models were tested in lab environments with lowered security guardrails, the government research body noted that, in a first for live testing, agents engaged in “social engineering” to manipulate human reviewers while pursuing unsanctioned tasks.
“This marks the first time AISI has documented unprompted, high-severity deception targeting a real person in a real-world setting,” the institute stated Tuesday, adding that no actual harm resulted from the incident.
The incident adds to a growing series of cases where advanced AI models have taken unauthorized actions fueling calls for stricter government regulation and a slower pace of AI development. While OpenAI and Anthropic previously disclosed late-July incidents where models escaped sandboxed environments to hack external systems, the British institute’s evaluation differed in one key way: researchers had deliberately granted the models live internet access during testing.
Across 122 cybersecurity challenges conducted by the institute, AI agents took “autonomous, unsanctioned action on the live internet, targeting real people and organizations” in 10 instances. The majority of these violations stemmed from Anthropic’s Mythos 5 model, with the remainder attributed to OpenAI’s GPT-5.6-Sol.
In the most severe case, an AI agent created multiple fake identities in an attempt to manipulate human reviewers into approving malicious code for an open-source project.
The AI agent directly contacted real individuals, using an online file-sharing service to send messages and files designed to trick people or their automated coding tools into executing malicious code, the institute reported. When its actions were flagged, the agent attempted to cover its tracks by altering earlier logs and considering a fresh identity to persist with the task.
The institute’s disclosure coincided with a White House meeting between leading AI executives and administration officials to discuss a new regulatory framework requiring government safety reviews before advanced AI models can be publicly deployed.
See also: Protect Yourself From AI Scams Key Online Safety Tips
In a statement posted on X, Anthropic noted that its models were evaluated under “deliberately permissive conditions,” which included the removal of standard safety filters and unrestricted internet access.
“We’re working closely with them to gather more details of the incident as we conduct our own investigation,” the company stated, adding that there was no indication the model broke out of a secure environment.
OpenAI categorized the two unauthorized behaviors as crossing outside the test environment and performing actions unnecessary for the assigned exercises.
“We are committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely,” OpenAI said in a Tuesday blog post.





