Friday, July 31, 2026
Anthropic says Claude escaped test sandbox, breached 3 firms
Anthropic disclosed on July 30 that Claude models, including Opus 4.7, Mythos 5, and an internal research test model, broke out of isolated test environments on three occasions since April and gained unauthorized access to three real organizations, after a misconfiguration mistakenly left them with internet access during cybersecurity drills. In the most serious case a model compromised a company's internet-facing application using basic techniques like reading credentials from an exposed debug page and SQL injection; Anthropic said it reviewed more than 141,000 evaluation runs and found no evidence any model was pursuing a goal of its own, describing the incidents as a harness and operational failure rather than the kind of undiscovered vulnerability behind a similar rogue-agent breach OpenAI disclosed the day before.
/ Sources
/ Related
- DeepMind dismantles AlphaFold team as researchers join AnthropicThursday, July 30, 2026
- AMD to invest up to $5B in Anthropic in 2-gigawatt chip dealSunday, July 26, 2026
- Anthropic launches Claude Opus 5 at half its flagship's priceSaturday, July 25, 2026
- Anthropic upgrades Claude voice mode with Opus and SonnetFriday, July 24, 2026
