Model Overview This repository hosts an NVFP4 quantized version of the Qwen3.6 27B model. The quantization process was executed using llm compressor , employing a mixed precision strategy to drastically reduce memory footprint while preserving essential model capabilities. Deployment & Inference This model is highly optimized for local inference on high end consumer hardware. Local testing and evaluations were conducted under the following environment: Hardware: 1x NVIDIA RTX 5090 Inference Engine: vLLM KV Cache: FP8 Quantization Details To achieve optimal performance, we applied specific quantization configurations across the model's architecture, heavily supported by advanced modifiers: Quantized to NVFP4: Full attention layers, linear attention layers, and the MLP blocks. Retained in BF16 (Untouched): Vision components, MTP, lm head, and embeddings. Enhancements: Utilized SmoothQuant modifiers alongside GPTQ modifiers to improve the overall post quantization performance. Calibration Configuration Customized calibration dataset with 512 samples and each 8192 sequence length. Evaluation & Benchmarks Benchmark BF16 (Alibaba Cloud) This Model (Local RTX 5090) Delta : : : : : : : MML…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy