Qwen3.6 27B NVFP4 GGUF GGUF conversion of nvidia/Qwen3.6 27B NVFP4 for use with llama.cpp . This is an NVFP4 quantized standalone language model. The converted GGUF preserves the NVFP4 tensors and the model's native MTP tensors. It is not a separate draft model and does not require an external speculator. Model Details Source model: nvidia/Qwen3.6 27B NVFP4 Base model: Qwen/Qwen3.6 27B Format: GGUF Quantization: NVIDIA NVFP4 / ModelOpt Architecture: Qwen3.6 27B dense model Purpose: Local inference and native MTP speculative decoding with llama.cpp NVIDIA quantized the weights and activations of linear operators inside the transformer blocks to NVFP4. Other tensors may remain in higher precision formats. This repository contains the complete target model. It is not an MTP, EAGLE3, or DFlash draft only checkpoint. Compatibility A recent version of llama.cpp with Qwen3.6, NVFP4, and native MTP support is required. Tested with: Windows NVIDIA GeForce RTX 5070 Ti 16 GB NVIDIA GeForce RTX 5060 Ti 16 GB llama.cpp CUDA backend Native Qwen3.6 MTP speculative decoding Older llama.cpp builds may fail to recognize the nvfp4 tensor type or may not correctly load the associated scale tensors. Pe…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy