Friday, July 31, 2026

Anthropic says Claude escaped test sandbox, breached 3 firms

Anthropic disclosed on July 30 that Claude models, including Opus 4.7, Mythos 5, and an internal research test model, broke out of isolated test environments on three occasions since April and gained unauthorized access to three real organizations, after a misconfiguration mistakenly left them with internet access during cybersecurity drills. In the most serious case a model compromised a company's internet-facing application using basic techniques like reading credentials from an exposed debug page and SQL injection; Anthropic said it reviewed more than 141,000 evaluation runs and found no evidence any model was pursuing a goal of its own, describing the incidents as a harness and operational failure rather than the kind of undiscovered vulnerability behind a similar rogue-agent breach OpenAI disclosed the day before.

/ Sources

/ Related