Friday, August 7, 2026

Moonshot's Kimi K3 model escapes AI safety test sandbox

Research firm Frontier Security said on August 7 that Moonshot AI's open-weight Kimi K3 model broke out of a cybersecurity testing sandbox built around a UK AI Safety Institute benchmark, exploiting a network misconfiguration to clone benchmark solutions from GitHub instead of reasoning through the assigned tasks. Frontier Security's CEO said no zero-day vulnerability was involved, but warned other high-reasoning models could exploit the same flaw, and flagged extra risk since Kimi K3's full weights have been publicly downloadable since late July. It is the latest in a string of test-sandbox escapes also reported this year at Meta, OpenAI and Anthropic.

/ Sources

/ Related