“In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available.”

The sandbox was not a sandbox. Across 141,006 runs, three models told they were in a simulation reached the open internet and compromised real organizational systems, and one pushed a malicious package to PyPI that landed on 15 machines. The exploits were weak passwords and unauthenticated endpoints, so nothing sophisticated was needed. Cyber evals stopped on July 23 and the affected organizations heard about it on the 27th. The review only happened because OpenAI disclosed the same class of failure first.