gdubicki/Qwen3 Coder Next NVFP4 GB10 Public mirror of saricles/Qwen3 Coder Next NVFP4 GB10 . Weights are byte identical to the upstream quant ( config.json and model.safetensors.index.json SHA 256 verified). This mirror exists to provide a pinned, stable, ungated reference for the qwen3 coder next deployment project on DGX Spark (GB10). Use the upstream repo if you want to track author updates. Credits Base model: Qwen/Qwen3 Coder Next by Alibaba / Qwen team (Apache 2.0) NVFP4 quantization: saricles using LLM Compressor with LLMCOMPRESSOR MOE CALIBRATE ALL EXPERTS=1 (all 512 experts calibrated) Calibration data: HuggingFaceH4/ultrachat 200k (64 samples × 2048 tok) License: Apache 2.0 (inherited from base model; redistribution permitted) Model details Architecture: qwen3 next — Hybrid DeltaNet linear attention + full attention + latent MoE Layers: 48 total (36 DeltaNet linear attention, 12 full attention) Parameters: 79.7B total, ~3B active per token (512 experts, 10 active + 1 shared) Quantization: NVFP4 via compressed tensors ; lm head , embed tokens , linear attn layers, mlp.gate , mlp.shared expert gate kept in BF16 Size on disk: 45.9 GB (70% reduction from ~149 GB BF16) KV cach…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy