“Pure LLMs score 0% on ARC-AGI-2, and public AI reasoning systems achieve only single-digit percentage scores.”
Pure LLMs score 0% on ARC-AGI-2 and the best reasoning systems score in single digits. At least two humans solved every task in two tries. o3 spent about $200 per task to reach 4%. The benchmark now counts cost.