“We found that LRMs have limitations in exact computation: they fail to use explicit algorithms and reason inconsistently across puzzles.”
Apple tested reasoning models on puzzles with adjustable difficulty, and every one hit zero accuracy past a threshold. Near the collapse they used fewer reasoning tokens, with budget to spare. Handing them the solution algorithm did not help. Apple is behind in AI, and this paper says the leaders’ flagship feature is weaker than advertised.