Cosmos Reason2 8B NVFP4 NVFP4 quantized version of nvidia/Cosmos Reason2 8B by vrfai using llm compressor. License: This model inherits the NVIDIA Open Model License from the base model. Commercial use and derivative models are permitted under its terms. NVFP4 Quantization Details Base model nvidia/Cosmos Reason2 8B Quantization NVFP4 — weights FP4, activations FP4 (dynamic local), scales FP8 Format compressed tensors (native vLLM support) Tool vllm project/llm compressor Model size → (~58% reduction) Requires NVIDIA Blackwell GPU (SM 120+), vLLM ≥ 0.19 What's Quantized / What's Not Component Precision Reason All LLM layers — FFN + attention projections (36 layers) NVFP4 Standard transformer, stable under 4 bit Vision encoder — all 27 blocks + merger BF16 Preserved for visual perception quality DeepStack merger list (3×) BF16 Multi scale visual fusion, sensitive to precision lm head BF16 Output logits preserved for generation stability Quantization Config (llm compressor) Quick Start (vLLM) Python (Transformers) OpenAI compatible API Tested Environment Component Version vLLM 0.19.1 Transformers 5.6.0 PyTorch 2.10.0+cu128 CUDA 12.8 (nvcc 12.8.61) llm compressor compressed tensors 0.…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy