◆ Rogue Quants · NVFP4 🪐 Qwopus3.6 27B Coder · NVFP4 💻 Coder 27B coder vision language · agentic + tool calling · thinking · GPTQ NVFP4 W4A4 ⚙️ NVFP4 · W4A4 💾 ~18 GB 📉 PPL 6.63 📐 256K context 🚀 vLLM · Blackwell 💻 Coder 🛠️ Tool calling Size on disk 18 GB vs 55.6 GB bf16 (~33%) wikitext 2 PPL 6.63 near lossless vs bf16 Context 256K 262144 tokens Scheme NVFP4 W4A4 · GPTQ + MSE TL;DR: Qwopus3.6 27B Coder, quantized to NVFP4 (W4A4) for vLLM on NVIDIA Blackwell. 18 GB, wikitext 2 PPL 6.63, 256K agentic coder. Qwopus3.6 27B Coder NVFP4 NVFP4 (W4A4) quantization of Jackrong/Qwopus3.6 27B Coder, packed in the compressed tensors nvfp4 pack quantized format with llm compressor. Weights are quantized with GPTQ (error compensated rounding) and an MSE observer, on a domain matched calibration blend that includes code. Near lossless. Fused layers (q/k/v, gate/up) share one NVFP4 global scale, so vLLM loads it cleanly with no per layer scale warning or fallback. wikitext 2 perplexity for this build: 6.63. About 18 GB on disk versus about 55.6 GB for the bf16 source (about 33%). Built for vLLM on NVIDIA Blackwell, where both the 4 bit weight and 4 b…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy