Alibaba's Qwen Grew From 7B to 2.4 Trillion Parameters in Three and a Half Years
Alibaba Cloud's Qwen family has grown from a 7-billion-parameter open-weight release in August 2023 to open weights with 2.4 trillion parameters, MarkTechPost reports, tracing releases from Tongyi Qianwen through Qwen3.
Alibaba Cloud began handing invitation codes for Tongyi Qianwen to corporate customers on April 7, 2023, four months after ChatGPT's release. The name draws partly on the philosopher Mencius. Four days later, at the Alibaba Cloud Summit in Beijing, then-CEO Daniel Zhang unveiled the model publicly, and Alibaba said it would fold it into every business, starting with DingTalk and the Tmall Genie voice assistant, China Daily reported.
The open-source turn came on August 3, 2023, when Alibaba released Qwen-7B and Qwen-7B-Chat as a direct answer to Meta's Llama 2. Qwen-7B was pretrained on more than 2.2 trillion tokens with a 2,048-token context, and its license allowed free commercial use below 100 million monthly users. Qwen-VL, the first vision-language branch, followed in late August, and Tongyi Qianwen opened to the general public on September 13, 2023, a sign of Chinese regulatory approval. The team published the Qwen Technical Report on arXiv the same month and closed the year with 72B and 1.8B downloads around December 1.
The year 2024 brought three generations in eight months. Qwen1.5 arrived on February 5 as a beta of Qwen2, covering dense models from 0.5B to 72B with stable 32K context at every size, plus a 14B mixture-of-experts model with 2.7B active parameters and a 110B dense model that was the family's first above 100B. Qwen2 followed in early June in five sizes from 0.5B to 72B, including Qwen2-57B-A14B, its first open mixture-of-experts flagship. Training data added 27 languages beyond English and Chinese, the 7B and 72B instruct models handled up to 128K tokens, and every size except 72B moved to the Apache 2.0 license.
Qwen2-Math and Qwen2-Audio launched in August, Qwen2-VL at the end of that month could analyze videos longer than 20 minutes, and on September 19, 2024, at the Apsara Conference, Alibaba released more than 100 open-source models at once under the Qwen2.5 name. The Qwen2.5 Technical Report says pretraining data grew from 7 trillion to 18 trillion tokens, with sizes running from 0.5B to 72B, 128K context and 8K-token generation. Alibaba said the family had passed 40 million downloads and inspired more than 50,000 derivative models on Hugging Face. Qwen2.5-Coder shipped its full family on November 11, QwQ-32B-Preview arrived in late November under Apache 2.0 as an open challenger to OpenAI's o1, and QVQ-72B-Preview, an experimental visual reasoning model, closed the year on December 24.
January 2025 belonged to DeepSeek-R1, and Qwen responded within weeks. On January 26 the team shipped Qwen2.5-VL in 3B, 7B and 72B sizes; TechCrunch noted it could control PCs and phones. Three days later, on the first day of the Lunar New Year, Alibaba launched Qwen2.5-Max, a large-scale mixture-of-experts model, and Reuters reported Alibaba's claim that it surpassed DeepSeek-V3. On March 6, QwQ-32B, built on Qwen2.5-32B and trained with reinforcement learning, shipped under Apache 2.0; Alibaba claimed performance comparable to DeepSeek-R1, a 671B model, and VentureBeat put the hardware gap at about 24 GB of VRAM against more than 1,500 GB. Alibaba's Hong Kong shares rose more than 7% that day. Qwen2.5-Omni-7B, released on March 26 under Apache 2.0, takes text, images, audio and video as input and answers in text or speech; CNBC framed it as a model for cost-effective AI agents.
Qwen3 arrived on April 29, 2025, Beijing time, with six dense models from 0.6B to 32B and two mixture-of-experts models, led by Qwen3-235B-A22B with 22B active parameters, all under Apache 2.0. The defining feature was hybrid thinking, letting a single model reason step by step or answer instantly at the user's toggle. The Qwen3 Technical Report lists pretraining on about 36 trillion tokens, and language coverage jumped from 29 to 119 languages and dialects; TechCrunch called it a family of hybrid reasoning models. On July 21, 2025, Qwen separated hybrid thinking again: Qwen3-235B-A22B-Instruct-2507 shipped as a non-thinking model with a 262K native context, and separate Thinking-2507 checkpoints followed.