Ornith 1.0 35B NVFP4 NVFP4 post training quantization of deepreinforce ai/Ornith 1.0 35B (Qwen3.5 MoE, 34.7B params) produced with the NVIDIA Model Optimizer. Quantized, validated and published by Robert Ressl on a single NVIDIA RTX PRO 6000 Blackwell Workstation Edition, using nvidia modelopt 0.44.0. Status This checkpoint is in the Qwen3 5MoeForConditionalGeneration / qwen3 5 moe form (the same form as nvidia/Qwen3.6 35B A3B NVFP4 ) and is validated to load and generate on vLLM (nightly, quantization modelopt , attention backend flashinfer , moe backend marlin ) on a Blackwell GPU. Quantization Format: NVFP4 (4 bit floating point, block size 16) — experts only quantization (MoE expert weights quantized to NVFP4 via ModelOpt's NVFP4 EXPERTS ONLY CFG ; attention QKV projections, shared experts and the vision encoder kept in higher precision for accuracy). Tool: nvidia modelopt 0.44.0 — mtq.quantize + export hf checkpoint (Unified HF checkpoint). Calibration: 512 samples from cnn dailymail , seq len 512, max algorithm. The full Qwen3 5MoeForConditionalGeneration model is quantized (language model experts NVFP4, vision encoder in BF16), matching the modelopt VLM flow. Size: ~23 GB (v…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy