Qwen3 Next 80B A3B Instruct FP8 Over the past few months, we have observed increasingly clear trends toward scaling both total parameters and context lengths in the pursuit of more powerful and agentic artificial intelligence (AI). We are excited to share our latest advancements in addressing these demands, centered on improving scaling efficiency through innovative model architecture. We call this next generation foundation models Qwen3 Next . [!Note] This repository contains the FP8 quantized Qwen3 Next 80B A3B Instruct model checkpoint for convenience and performance. The quantization method is "fine grained fp8" quantization with block size of 128. You can find more details in the quantization config field in config.json . In addition, the experimental results presented in this model card are obtained from the original bfloat16 model prior to FP8 quantization. Highlights Qwen3 Next 80B A3B FP8 is the first installment in the Qwen3 Next series and features the following key enchancements: Hybrid Attention : Replaces standard attention with the combination of Gated DeltaNet and Gated Attention , enabling efficient context modeling for ultra long context length. High Sparsity Mixt…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy