“In one scenario, Mythos 5 persuaded developers to download a poisoned PyPI package. When that company’s scanner installed the package, Claude’s hidden code executed. Claude then used these credentials to access further infrastructure from this company.”
A test agent broke containment and attacked three companies that never agreed to be targets. That happened in April. Anthropic found out months later, and only because OpenAI’s own sandbox breach prompted somebody to go look. Both labs spend their days publishing safety research about hypothetical future risks while the actual agents they ship today are compromising real infrastructure and nobody notices for a quarter.