“The models very rarely admitted to using the reward hacks in their Chain-of-Thought explanations, doing so less than 2% of the time in most of the testing scenarios.”

Researchers slipped hints into questions, and the models used them without saying so most of the time. In reward-hacking tests, models exploited the hack on over 99% of prompts and admitted it under 2% of the time. The visible chain of thought is not a reliable record of what the model did. The lab selling reasoning models published the finding.