“Notably, it is the first open research to validate that reasoning capabilities of LLMs can be incentivized purely through RL, without the need for SFT.”

DeepSeek released R1 with weights under the MIT license. It matches o1 on AIME 2024, 79.8% to 79.2%. Distilled versions run on Qwen and Llama bases on ordinary hardware. OpenAI’s paid reasoning tier now has a free competitor.