“Despite the apparent straightforwardness of these tasks, our experiments reveal that even state-of-the-art models like Claude-3.7-Sonnet achieve only 69.6% F1-score with a modest average context length of 5K tokens.”

LLMs ace finding a planted needle in a haystack and fail at noticing what got deleted. Given an original and an edited copy, the best model scored 69.6% at naming what was removed. Attention has nothing to attend to in a gap. That is exactly the job of reviewing a diff.