“While Qwen2.5 was pre-trained on 18 trillion tokens, Qwen3 uses nearly twice that amount, with approximately 36 trillion tokens covering 119 languages and dialects.”
Alibaba released Qwen3 under Apache 2.0, trained on 36 trillion tokens. It claims parity with o1 and Gemini 2.5 Pro on its own benchmarks. Part of the training data came from its own earlier models. Open weights keep closing the gap.