Qwen3.8-Omni-Flash Explained

As of 2026-09-26, Alibaba's Qwen team shipped Qwen3.8-Omni-Flash on 2026-09-18: native omni (text, image, audio, and video in; text out) with a 1M context window. If you need chapters from a meeting or a long video, paste the link — you do not have to wire an omni API.

Shipped · 2026-09-18 Native omni · 1M context Live demo below

Get the notes, skip the omni API

Paste a long-form video or meeting URL — BibiGPT turns it into chapters, a transcript, and follow-up Q&A.

Add BibiGPT as a preferred source on Google See more BibiGPT in Top Stories and AI answers.

Key facts (90-second read)

As of 2026-09-26, Alibaba's Qwen team shipped Qwen3.8-Omni-Flash on 2026-09-18: native omni (text/image/audio/video → text), a 1M context window, and a reported +26% lift versus Qwen3.5-Omni-Plus on 30 evals. Last updated 2026-09-26.

Features

What Qwen3.8-Omni-Flash actually is

On 2026-09-18 the Qwen team shipped Qwen3.8-Omni-Flash as a native omni model: one stack that takes text, image, audio, and video and returns text. It is live on Qwen AI and Alibaba Cloud Model Studio. This is not the 2026-08-26 text Flash-Next checkpoint.

Native omni: four inputs, text out

Qwen describes a single model that accepts text, image, audio, and video and writes text. That is a different product from Qwen3.8-Flash (text multimodal MoE) and from Qwen3-ASR (speech recognition only).

1M context, 131K max output

The documented context window is one million tokens, with a 131K max output. A long meeting transcript can fit in one call. A wide window still does not file chapters or timestamps for you.

Vendor evals vs Qwen3.5-Omni-Plus

Qwen reports a +26% average lift versus Qwen3.5-Omni-Plus across 30 evals, and large drops on AliMeeting DER / cpWER (88.11 / 89.61 → 3.35 / 17.18). Those are vendor scores, not a BibiGPT benchmark.

What this means if you work with meetings and long video

A cheaper omni API helps teams that already wire Alibaba Cloud. It does not, by itself, give a student timestamped notes from a two-hour Bilibili upload or a recorded standup. BibiGPT is the product for that second job.

You should not have to wire DashScope yourself

Paste a link, get chapters, a transcript, and follow-up Q&A. BibiGPT supports Qwen3.8 Omni Flash for long-form notes — you do not pick an omni recipe.

Omni input still needs a notebook

Audio and video in is useful. A 90-minute meeting still needs timestamps and an export. One cheap API call that you never file is wasted; a chaptered summary is a document you can search next week.

This is not Qwen3.8-Flash and not Gemini 3.8 Flash

Qwen3.8-Flash is the August text / multimodal MoE. Gemini 3.8 Flash is Google's same-week competitor on AV tasks. Qwen3-ASR is speech recognition. This page is the 2026-09-18 Omni-Flash event. One intent, one URL.

5 key changes (90-second read)

Headline facts from Qwen's 2026-09-18 Qwen3.8-Omni-Flash release.

  1. 1

    Shipped 2026-09-18 on Qwen AI and Model Studio

    Qwen released Qwen3.8-Omni-Flash the same day on Qwen AI and Alibaba Cloud Model Studio (DashScope id qwen3.8-omni-flash). It is a hosted omni SKU, not an open-weight Flash-Next drop.

  2. 2

    Native omni: text, image, audio, video → text

    One model accepts four input types and writes text. That is a different job from Qwen3.8-Flash (text multimodal MoE) and from Qwen3-ASR (speech recognition).

  3. 3

    1M context, 131K max output

    Qwen documents a one-million-token window. LiteLLM's day-0 note lists 131K max output. Long meetings fit in one call; they still need chapters after the call.

  4. 4

    Vendor evals vs Qwen3.5-Omni-Plus and Gemini 3.8 Flash

    Qwen reports +26% average on 30 evals versus Qwen3.5-Omni-Plus, AliMeeting DER/cpWER 88.11/89.61 → 3.35/17.18, and AV scores close to Gemini 3.8 Flash. Vendor numbers, restated here.

  5. 5

    International list price $0.15 / $0.47 per million tokens

    LiteLLM listed $0.15 input, $0.016 cached, $0.47 output per million tokens, and video input about 89% cheaper than Qwen3.5-Omni-Plus. Confirm on Model Studio before you budget.

3 typical scenarios for BibiGPT users

Where a native omni Qwen SKU matters — and where a summarizer is the actual product.

You already call DashScope

Use Qwen3.8-Omni-Flash in your own stack for omni understanding. Then still paste the finished lecture or standup into BibiGPT so humans get chapters instead of a raw token dump.

You just need notes from a long Chinese video or meeting

A two-hour Bilibili upload or a recorded standup does not need an omni API brief. Paste the URL, export the transcript, keep asking questions. That is the BibiGPT path.

You are comparing Qwen SKUs this month

Qwen3.8-Flash is the August text MoE. Qwen3-ASR is speech recognition. Gemini 3.8 Flash is the Google AV competitor. This page is Omni-Flash only. One model-family event, one URL.

Related BibiGPT tools

Meeting and long-video notes workflows that pair with this release.

Sources

Specs on this page come from Qwen's 2026-09-18 announcement and LiteLLM's same-week day-0 note.

  • Qwen shipped Qwen3.8-Omni-Flash on 2026-09-18 as a native omni model (text/image/audio/video → text) with a 1M context window, live on Qwen AI and Alibaba Cloud Model Studio. Qwen reports +26% versus Qwen3.5-Omni-Plus on 30 evals and AliMeeting DER/cpWER 88.11/89.61 → 3.35/17.18.

    Qwen — Qwen3.8-Omni-Flash ↗
  • LiteLLM's 2026-09-18 day-0 note lists DashScope id qwen3.8-omni-flash, 1M context, 131K max output, international list prices of $0.15 input / $0.016 cached / $0.47 output per million tokens, and video input about 89% cheaper than Qwen3.5-Omni-Plus.

    LiteLLM — Qwen3.8-Omni-Flash day-0 ↗

Terms used on this page

What is a native omni model?

A native omni model takes more than one media type in the same stack — here text, image, audio, and video — and writes a single text reply. It is not a cascade of a separate ASR model plus a text LLM. Qwen3.8-Omni-Flash is documented as that kind of stack.

What is a 1M context window good for in meeting notes?

A one-million-token window can hold a long transcript plus instructions in one call. It does not, by itself, produce chapters, timestamps, or an export you can search next week. Those artifacts are the job of a notes product.

Loved by creators, students & researchers

Why people use BibiGPT to turn videos into text every day.

Trusted by 50,000+ users worldwide

★★★★★

“I paste a link and get clean captions in seconds — it saves me hours of retyping every single week.”

Maya R.

Content Creator · Repurposes short videos

★★★★★

“Exporting the transcript lets me review new words at my own pace instead of pausing the video constantly.”

Daniel K.

Language Learner · Studies with real videos

★★★★★

“Accurate, timestamped text I can quote directly. It has quietly become part of my daily workflow.”

Priya S.

Researcher · Cites public talks

Frequently Asked Questions

Ask us anything!

Popular guides

Skip the omni API. Get the notes.

Paste a YouTube, Bilibili, podcast, or meeting link. BibiGPT returns chapters, a transcript, and follow-up Q&A. Last updated 2026-09-26.