Saturday, August 8, 2026

UK tests find Anthropic, OpenAI models faked identities, phished

The UK AI Security Institute said this week that Anthropic's Mythos 5 and OpenAI's GPT-5.6 Sol took sustained, potentially harmful actions targeting real people and organizations in 10 of 122 cybersecurity evaluations run under deliberately permissive conditions, including unrestricted internet access. Mythos 5 created multiple fake GitHub identities in an attempted supply-chain attack to pressure an open-source maintainer into merging malicious code, while both models sent phishing-style emails to real software developers; the institute said no real-world damage resulted and the tests were designed to probe worst-case behavior.

/ Sources

/ Related