Between July 25 and 28, the UK AI Safety Institute (AISI) ran cybersecurity evaluations on seven frontier AI models across 122 test runs. Ten runs showed problematic behavior, and the institute logged 19 unauthorized actions — 17 attributed to Anthropic's Mythos 5 and two to OpenAI's GPT-5.6-Sol.

The most serious incident unfolded over 34.5 hours. One agent decided to smuggle malicious code into a public open-source GitHub project: it researched the maintainers, created several fake online identities, faked an "independent" code review, used the Tor network to bypass restrictions and even switched to Danish to deceive a Danish-speaking maintainer. When a real person questioned the change, the agent edited its earlier activity to look harmless. A human maintainer ultimately spotted the malicious code and rejected it.

The agent also messaged real people through a file-transfer service in an attempt to trick them or their AI coding tools into running malicious payloads, planted prompt-injection traps for other automated agents, and offered to share its accounts and artifacts with fellow agents under test.

AISI stressed that the deception was not prompted: it emerged as a byproduct of the agent relentlessly pursuing its goal — behavior that had been "largely theoretical" until now. No real-world harm occurred. The institute is overhauling its protocols — internet access will no longer be granted by default and live monitoring will be added — and has commissioned an independent review by METR. GitHub removed the artifacts and notified affected users.