Alibaba Opens Qwen3.8-Flash-Next, a 125B Preview of the Qwen4 Architecture
On August 26, 2026, Alibaba’s Qwen team open-sourced Qwen3.8-Flash-Next—a 125B multimodal MoE with 6B active parameters, 262K native context, and vendor scores that preview the Qwen4 architecture—at $0.16/$0.47 per million tokens on the production Flash API.
TLDR
Alibaba Qwen on Wednesday, August 26, 2026 (@Alibaba_Qwen, 12:34 UTC) released Qwen3.8-Flash-Next weights on Hugging Face / ModelScope. Architecture: 125B backbone + 51B N-gram embeddings + 4B MTP; 6B active per token; 262K native context, 1M with YaRN. Vendor benches: DeepSWE 1.1 58.7, SWE-bench Pro 62.5, CoWorkBench 73.9, LiveCodeBench v6 91.9, GPQA Diamond 91.7. Train cost claimed 1/9 of Qwen3.7-Plus. Production API Qwen3.8-Flash: $0.16 / $0.47 per 1M in/out. Qwen frames it as a Qwen4 architecture preview.
What shipped
| Item | Qwen Aug 26 |
|---|---|
| SKU | Qwen3.8-Flash-Next (open weights) |
| Active | 6B / token |
| Context | 262K / 1M YaRN |
| API | $0.16 / $0.47 |
Product-line de-dupe: not Aug 14 Qwen3.8-27B, not Jul 19 Qwen3.8-Max-Preview. New architecture preview, not a dual-date of 27B.
Independent Arena / Artificial Analysis numbers were not confirmed in this pass—label vendor scores as lab-claimed.
Why this story matters
A 6B-active 125B MoE at $0.16 input is Alibaba’s bid to make Qwen4 look cheap before the next closed Max. Watch: whether Flash-Next actually appears on LMArena, or stays a ModelScope drop.