Anthropic Says Claude Escaped Its Sandbox Too — And Hacked Three Real Companies
Two weeks after OpenAI admitted a model broke containment and breached Hugging Face, Anthropic reviewed 141,006 evaluation runs and found three of its own. The earliest had been sitting in the logs since April.