Ornith 1.0 35B NVFP4 NVFP4 (W4A4) quantization of deepreinforce ai/Ornith 1.0 35B — DeepReinforce's self scaffolding agentic coding model ( qwen3 5 moe , 35B MoE with a Qwen3 VL vision tower). Quantized with llm compressor to compressed tensors nvfp4 pack quantized . 21.9 GB (from 70.3 GB bf16). Serves on a pair of 16 GB GPUs. Loads in vLLM with no quantization flag (auto detected). What was quantized All linear layers → NVFP4 (W4A4, group size 16). Kept in bf16: the vision tower ( re:. visual. ), the MoE routers ( mlp.gate , mlp.shared expert gate ), and lm head . The 30,720 routed expert projections (256 experts × 3 × 40 layers) are per expert pack quantized. Benchmarks pass@1 on HumanEval+ / MBPP+, scored with an identical local harness. Quantized (this model) vs. a panel of same class open baselines: Benchmark no think think HumanEval+ (N=163) 87.1% 93.9% MBPP+ (N=160) 78.1% 80.6% With reasoning enabled, the W4A4 quant matches or tops the strongest same class open coders we benchmarked against, on both suites. Quality of the W4A4 quantization is intact. Reasoning model eval tip: Ornith reasons at length. For one shot code benchmarks (a) give it room ( max tokens ≥ 6500 ), and (…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy