“They used 16 NVIDIA H100s for 26 minutes per training run, that equates to around $6.”
S1 trained for about $6 a run on 16 H100s. The trick is 1,000 carefully chosen examples and a Wait token that forces longer thinking. The $6 sits on top of Qwen2.5, which someone else paid to train. The moat around reasoning models just got shallower anyway.