NVDA 230.36 ▲0.84%GOOGL 338.46 ▼1.17%MSFT 499.70 ▼2.04%AMD 477.57 ▲4.69%INTC 95.80 ▲4.51%TSMC 428.91 ▲2.85%AMZN 258.51 ▼0.15%META 616.77 ▲1.00%AAPL 319.97 ▼2.51%PLTR 174.33 ▼4.49%
Markets at last close

Alibaba · Models

Qwen previews Qwen4 architecture with Qwen3.8-Flash-Next

·1 min read

Qwen3.8-Flash-Next is an experimental open-weight release on Hugging Face that previews the architecture planned to underpin Qwen4. The model is distributed in Hugging Face Transformers format and is compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed and related serving stacks. Qwen describes the release as a cost-efficiency push for frontier foundation models as parameter counts and context windows continue to grow.

The architecture introduces Hybrid Attention with Qwen Sparse Attention, Gated Residual, N-gram Embedding and a tailored training recipe using Muon and AdamW for specific weight categories. The model is listed as a Causal Language Model with Vision Encoder with 125B parameters, 6B activated, 51B n-gram embedding parameters and 4B MTP. It supports 262,144 tokens natively and is extensible up to 1,000,000 tokens.

Reported benchmark results include 91.7 on GPQA Diamond, 62.5 on SWE-bench Pro, 81.0 on SWE-bench Multilingual and 64.4 Pass@3 on ClawEval-MM. Qwen recommends API-based integration for streamlined use, with dedicated serving engines such as SGLang, KTransformers or vLLM for production workloads and high-throughput scenarios.

Originally reported by huggingface.coRead the source →
Related coverage
All Alibaba news →