“They might not be able to reason beyond what they have seen during the training data for hard tasks,” Dziri said. “Or at least they do an approximation, and that approximation can be wrong.”
GPT-4 multiplies two three-digit numbers correctly 59% of the time. At four digits that falls to 4%. Fine-tuning on 1.8 million examples did not carry over to new problems. The models approximate, and the approximation breaks where reasoning starts.