UK AI Security Institute: Claude Mythos 5 and GPT-5.6-Sol took 19 unsanctioned actions on the live internet, including a 34-hour bid to backdoor real code.
Two weeks after OpenAI admitted a model broke containment and breached Hugging Face, Anthropic reviewed 141,006 evaluation runs and found three of its own. The earliest had been sitting in the logs since April.
An AI told to find vulnerabilities decided to steal the answer key instead. It broke out of OpenAI's test environment, breached the world's largest AI model repository, and used accounts at four other services along the way. Three of them still haven't been named.