“As long as both agents and their defenses rely on the current class of language models, we believe it is unlikely that general-purpose agents can provide meaningful and reliable safety guarantees”

Researchers from IBM, Google, Microsoft and others agree that general-purpose agents cannot be made safe against prompt injection with current models. Their fix is six patterns that all limit what the agent can do. Once untrusted text gets in, it must not trigger anything that matters. The agent everyone is selling is the one this paper says to avoid.