“This advancement stems from enhanced thinking depth during the reasoning process: in the AIME test set, the previous model used an average of 12K tokens per question, whereas the new version averages 23K tokens per question.”
DeepSeek’s AIME score jumped from 70% to 87.5%, and its tokens per question roughly doubled. A lot of that gain is the model thinking longer on the bill. The card claims fewer hallucinations while its own SimpleQA score dropped from 30.1 to 27.8. DeepSeek calls it a minor version upgrade and ships it under MIT anyway.