“It is not proof of a distillation, but it does show that Claude’s self-concept is embedded in these models’ weights.”
MATS researchers Benji Berczi and Kyuhee Kim tested GLM 5.2 and Kimi K3. Tell GLM 5.2 it is Claude and its uncensored rate on sensitive PRC topics jumps from 17% to 85%. A prompt string should not move a safety-relevant behavior by 68 points. Either the training data carried the persona along with the censorship policy, or the alignment was never deeper than a system prompt.