Saturday, August 15, 2026

Anthropic raises misalignment risk rating, discloses Model 2

Anthropic published its second company-wide Risk Report on August 14, upgrading its assessed risk of catastrophic harm from misalignment in high-stakes settings from very low to low, citing growing uncertainty rather than a specific failed safety test. The report also disclosed an unreleased internal model called Model 2, said to be somewhat more capable than Anthropic's public frontier model Claude Mythos 5 but with no current plan for external release, and kept the risk rating for automated AI research and development at low while noting that Claude now writes a large majority of the code merged into Anthropic's own production systems.

/ Sources

/ Related