Alibaba Cloud listed Qwen3.8-Omni-Flash as a new model in its Model Studio release notes on September 17, 2026. It is an omni-modal model: it accepts text, images, audio and video in one request and returns text. Alibaba describes it as suited to multimodal analysis and agent applications, with thinking and non-thinking modes, function calling, web search, and context caching for audio and video understanding.
The model documentation gives a 1M-token context window, with maximum input of 991,808 tokens in non-thinking mode and 983,616 in thinking mode, and up to 131,072 output tokens. Thinking is on by default with adjustable reasoning effort, it supports spatial audio through a multichannel option, and it covers 113 languages and dialects. It is available in China (Beijing), Singapore, China (Hong Kong), Japan (Tokyo), Germany (Frankfurt) and US (Virginia) regions.
A million-token window matters more for audio and video than for text, because long recordings consume tokens quickly; a model that can hold long recordings in context and then call tools on them is aimed at the audio-visual agent workloads Alibaba is courting. It extends the Qwen3.8 generation, which already includes open-weight and Flash text releases, into a single model for every input type.
Unlike several Qwen3.8 siblings, the documentation fetched for this entry does not describe open weights; it is offered through Alibaba's cloud API. Performance claims about the model circulating in press coverage were not found in the Alibaba pages consulted here, so this entry makes no benchmark claims.