barryke/Ornith 1.0 9B FP8 DYNAMIC FP8 dynamic activation, per channel weight quantization of deepreinforce ai/Ornith 1.0 9B , produced with LLM Compressor and stored in the compressed tensors format that vLLM and Transformers ≥ 5.8.1 load directly. Base model deepreinforce ai/Ornith 1.0 9B (Qwen 3.5 9B, multimodal, reasoning) Quantization scheme FP8 DYNAMIC (E4M3 weights, E4M3 activations, dynamic per token scale) Weight granularity per channel Activation granularity dynamic per token Layers quantized all Linear layers except those listed below Layers kept at BF16 lm head , re:. visual. (vision tower), re:. linear attn. (hybrid Gated DeltaNet projections) Calibration data none — dynamic activations need no calibration set Framework llmcompressor==0.12.0 , compressed tensors==0.17.1 Quantization hardware NVIDIA H100 (via Modal) License MIT (inherited from the base model) Why this quantization? FP8 DYNAMIC is the simplest scheme that recovers near full accuracy on most LLMs while cutting weight size roughly in half: No calibration dataset required — activation scales are computed on the fly per token at inference time. Per channel weight scales avoid the accuracy loss that per tensor…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy