Qwen3.5 9B — GGUF (imatrix) weight quant set A complete set of importance matrix (imatrix) GGUF quantizations of Qwen/Qwen3.5 9B , spanning the full ladder from IQ1 up to Q8 0 , plus the BF16 baseline. This set exists as a canonical reference for cross quant / cross backend quality (perplexity) and throughput comparison. Every quant in this repository was produced from the same BF16 source and the same importance matrix ( Qwen3.5 9B imatrix.gguf , included here for full reproducibility). Generated with the jimbothigpen/llama.cpp fork. These GGUFs and their imatrix are produced by that fork's quantization toolchain. It is a fork of upstream ggml org/llama.cpp that integrates features ported from several llama.cpp forks — including the IQ K / IQ KS / IQ KT (trellis) quant families from ik llama.cpp — plus ROCm and Vulkan backend support. Provenance Base model Qwen/Qwen3.5 9B (unmodified upstream weights) GGUF source BF16 GGUF produced by convert hf to gguf.py … outtype bf16 Quantizer llama quantize from the jimbothigpen/llama.cpp fork (a fork of ggml org/llama.cpp) imatrix Qwen3.5 9B imatrix.gguf (included) — md5 c81ec1fbddcbda8d8b07644deb316460 These are derivative weights of Qwen/Q…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy