“A single prompt can be designed to work across all of the major frontier AI models.”
HiddenLayer found one prompt, dressed up as a config file and a TV script, that breaks every major model’s guardrails. It pulls out system prompts and dangerous instructions in leetspeak. HiddenLayer sells AI security, so it has a stake. The guardrails failed anyway.