AdvertisementAdvertisementAdvertisementAdvertisement
AI

Alibaba Introduces Multimodal Model Qwen3.8-Omni-Flash

9/22/2026, 08:14 AM • Evgenia Sliv

(edited: 09/22/2026)

Alibaba Introduces Multimodal Model Qwen3.8-Omni-Flash

Alibaba Cloud has introduced Qwen3.8-Omni-Flash – a multimodal AI model that simultaneously works with text, images, audio recordings, and video. The system's context window reaches 1 million tokens, allowing it to process large volumes of information within a single request. The main version generates a text response, and the model itself supports calling external functions, web search, and a reasoning mode with adjustable depth. Developers position Qwen3.8-Omni-Flash not only as a tool for recognizing and analyzing multimedia data but also as a foundation for performing complex sequential tasks using AI agents. Separately, Alibaba Cloud has introduced a version of Qwen3.8-Omni-Flash-Realtime, designed for real-time audio and video processing.

One of the main application areas for the model, according to Alibaba, is the processing of long audio and video recordings. Qwen3.8-Omni-Flash can analyze multi-hour materials, find the necessary episodes, and generate structured results. When working with meeting recordings, the system can create transcripts, distinguish between participants in the conversation, highlight tasks set, and prepare summary reports. For video, the user can predefine a time interval, search object, required level of detail, and format of the final response. This approach allows the use of a single model for different types of multimedia content without the need to separately convert the original data into text. Additional tools expand the system's capabilities in scenarios related to translation, video editing, creating comments, and other media material processing.

According to Alibaba Cloud, Qwen3.8-Omni-Flash showed on average more than 25% higher results compared to Qwen3.5-Omni-Plus in 29 internal and industry tests. The company also claims a reduction in the cost of processing audio and audiovisual content through the API relative to the previous model. Qwen3.8-Omni-Flash is already available through Alibaba Cloud Model Studio and supports Chat Completions and Responses interfaces compatible with the OpenAI API. Thus, developers can connect the model to existing applications and use it not only for individual requests but also as part of automated processes. A separate realtime version extends this approach to scenarios where audio and video analysis must occur directly during interaction.

Popular news