Evidence at a glance

18T token, 7TQwen2.5
100Qwen2.5
0.6B 235B parametersQwen3
119 , 29Qwen3
Qwen2.5-72B Llama 3 405BEvidence
2.4T parameters, 95B parametersQwen3.8

The mechanism in one line

InputReduce the input to a workable scale

Compress the visual or contextual input before the main reasoning path.

MechanismSpend compute where it matters

Route or verify the expensive step instead of repeating the full path.

OutcomeEnd with a measurable workflow result

Translate the mechanism into a bounded deployment or evaluation check.

The Starting Point Was Not Scale, but the Pace of Release

Alibaba Cloud introduced Tongyi Qianwen in April 2023, initially offering invited corporate customers access before planning to integrate the model into services such as DingTalk and the Tmall Genie. For engineering teams, this looked like a familiar cloud provider launch of a generative AI service. The trajectory changed more decisively in August, when Alibaba released the Qwen-7B and Qwen-7B-Chat weights. Only a few months separated an invite-only experience from downloadable models.

That move was more than the release of a chatbot. The supplied account says Qwen-7B was pretrained on more than 2.2 trillion tokens; a vision-language branch followed in August, and Tongyi Qianwen opened to the public in September. A pattern emerged that would recur: enter the market through a service or demonstration, then broaden access to weights, model sizes, and capabilities. For developers, the ability to obtain models and experiment locally became a defining part of the product line.

The Model Family Began to Branch by Task

By 2024, Qwen could no longer be described as a single model getting upgraded. Qwen1.5 ranged from 0.5B to 72B and offered a stable 32K context window. Qwen2 expanded to five sizes, including a 57B-A14B mixture-of-experts model. Qwen2.5 pushed further into coding and reasoning: the supplied timeline reports 18 trillion training tokens for the generation, followed late in 2024 by Qwen2.5-Coder, the QwQ-32B-Preview reasoning model, and the experimental visual-reasoning model QVQ-72B-Preview.

The shift is important: the family was no longer just one model made larger or smaller. It was branching beyond general conversation into coding, reasoning, vision, and multimodal work. In 2025, Qwen2.5-VL was described as able to operate computers and phones, while Qwen2.5-Omni-7B accepted text, images, audio, and video and returned text or speech. For model selection, a version number alone no longer says whether a model fits. Teams first need to identify whether they need coding, visual understanding, or input and output across modalities.

Total Parameters Do Not Equal the Scale You Must Run

Qwen’s later releases put efficiency alongside raw scale. The supplied timeline lists Qwen2 at 57B-A14B, Qwen3-Next at 80B-A3B, and Qwen3.5 at 397B-A17B. These labels describe models by both total parameters and a smaller number in the suffix. As the naming indicates, total size and the parameters active for a given computation are not the same thing. A 2.4T total parameter count should therefore not be read as meaning that every request computes over all 2.4 trillion parameters.

That distinction matters for deployment, but it does not make a large model automatically cheap or lightweight. Qwen3.8-2.4T-A95B is listed as an open-weight release, with 95B in its active-parameter designation; storage, memory, and weight-handling costs still need to be accounted for. Another path is a smaller, task-focused model: Qwen3.6-35B-A3B targets agentic coding, while Alibaba claimed that a dense 27B model outperformed Qwen3.5-397B on coding. That is a vendor claim about a particular capability, not a substitute for evaluation on a team’s own workload. It does, however, underline that a larger model does not automatically perform better on every task.

Open Weights Do Not Mean Uniform Licensing

It is tempting to summarize Qwen’s history as becoming steadily more open, but the licensing record in the supplied material does not support treating openness as a simple switch. The early Qwen-7B terms tied commercial use to monthly active-user scale. Apache 2.0 became common with Qwen1.5, and Qwen2 is listed as using it for every size except 72B. At the same time, flagship models such as Qwen2.5-Max and Qwen3-Max followed a closed or API-only path. Open weights were never a uniform policy across the entire portfolio.

The 2026 releases make model-by-model checking even more important. The account says Qwen3.8-2.4T-A95B requires an additional license above US$50 million in annual revenue, while Qwen3.8-27B uses Apache 2.0. Qwen-Image-2.1, meanwhile, is listed under a research-only license. In other words, “open weights” means that weights are available in some form; it does not guarantee identical rights for every use, business scale, or redistribution model. Product teams that look only at a download page or model name risk confusing technical access with commercial permission.

Choose Across Capability, Cost, and Constraints

For technical leaders, Qwen’s trajectory turns model selection from “pick the strongest model” into a portfolio decision. Teams need to compare general-purpose, coding, vision, and multimodal models by task; distinguish dense models from those labeled with active-parameter counts; and consider local deployment resources, API dependence, and licensing together. The examples in the timeline show that these choices do not always align: flagship reasoning can be offered through a closed service, smaller models can have open weights, and a very large model can still carry a revenue threshold.

Parameter figures also need to be separated from the strength of the evidence behind them. The material lists Qwen3.8-2.4T-A95B as an open-weight release in August 2026, but describes the 5–10T scales for Qwen 4.5 and Qwen 5 as projections, while saying Qwen 4 is still in training. A projection is not a specification for a released model. The actionable approach is to verify weights, license, and deployment requirements for each candidate, then test capability and cost on your own workload. Do not turn a roadmap, total parameter count, or a vendor’s single benchmark comparison into a procurement decision. Qwen’s evolution expands the range of choices, but it does not remove engineering costs or usage boundaries.