
Alibaba Group Holding Ltd. has released Qwen3.8-Flash — a reduced-price artificial intelligence model designed to accelerate the global adoption of the Qwen AI platform. The company has made the model available for download and published its weights — the parameters that help AI systems make decisions. According to Alibaba, Qwen3.8-Flash contains 125 billion parameters and is comparable in performance to the latest offerings from competitors, including Anthropic's Opus 4.6 and DeepSeek's V4-Flash.
The production version of the model is available via the QwenCloud API at $0.16 per million input tokens and $0.47 per million output tokens. The Qwen3.8-Flash architecture utilizes 125 billion parameters plus 51 billion N-gram embeddings, with only 6 billion of them activated per token, ensuring high energy efficiency when processing requests.
The model delivered the following results on industry benchmarks: 58.7 points on DeepSWE 1.1, 62.5 on SWE-bench Pro, 73.9 on CoWorkBench, 84.5 on AndroidWorld, and 95.7 on MathVision within confidence intervals. According to the company, Qwen3.8-Flash was trained at one-ninth the cost of the previous Qwen3.7-Plus model, while outperforming it on coding and office application tasks.
Alibaba also published the weights of Qwen3.8-Flash-Next — an early preview of the architecture planned for the next-generation Qwen4. This model natively supports a context of 262,144 tokens and is expandable to 1 million tokens using YaRN technology. The weights of both models are available on the Hugging Face and ModelScope platforms for research and commercial use.
The release of an optimized open-weights model reflects Alibaba's strategy of democratizing access to cutting-edge AI technologies and strengthening the Qwen platform's position in the global market. Reducing inference costs while maintaining competitive performance allows developers to integrate generative AI capabilities into their products with lower infrastructure expenses.
