“Just asking agents to ’test’ or repeatedly asking them to test more results in poor testing. It turns out that asking them to use test techniques also results in poor testing.”

Twenty-six testing approaches, from fuzzing to TLA+, and the plain default with no instructions beat most of them. The agents adopt the framework, write tests that look like tests, and verify nothing. Every vendor demo showing an agent “checking its own work” is showing you this. The tests pass because the tests are decorative.