NVDA 217.44 ▼1.51%GOOGL 335.02 ▼1.28%MSFT 501.02 ▼1.24%AMD 459.61 ▼2.36%INTC 88.97 ▼0.60%TSMC 414.00 ▼0.32%AMZN 254.92 ▼1.87%META 578.54 ▲1.08%AAPL 325.13 ▲2.61%PLTR 179.92 ▼3.47%
Markets at last close

Alibaba · Models

Qwen previews Qwen4 architecture with Qwen3.8-Flash-Next

·1 min read

Qwen3.8-Flash-Next is an experimental open-weight release on Hugging Face that previews the architecture planned to underpin Qwen4. The model is distributed in Hugging Face Transformers format and is compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed and related serving stacks. Qwen describes the release as a cost-efficiency push for frontier foundation models as parameter counts and context windows continue to grow.

The architecture introduces Hybrid Attention with Qwen Sparse Attention, Gated Residual, N-gram Embedding and a tailored training recipe using Muon and AdamW for specific weight categories. The model is listed as a Causal Language Model with Vision Encoder with 125B parameters, 6B activated, 51B n-gram embedding parameters and 4B MTP. It supports 262,144 tokens natively and is extensible up to 1,000,000 tokens.

Reported benchmark results include 91.7 on GPQA Diamond, 62.5 on SWE-bench Pro, 81.0 on SWE-bench Multilingual and 64.4 Pass@3 on ClawEval-MM. Qwen recommends API-based integration for streamlined use, with dedicated serving engines such as SGLang, KTransformers or vLLM for production workloads and high-throughput scenarios.

Originally reported by huggingface.coRead the source →
Related coverage
All Alibaba news →