Ornith 1.0 397B — MXFP4 (mixed precision) A 4 bit MXFP4 quantization of Ornith 1.0 397B, produced with qstream . Only the routed MoE experts (~95% of the weights) are quantized to MXFP4; everything quality sensitive — including the always on shared expert — stays BF16 , bit identical to the source. The original model card follows in full below. Size 225.9 GB (210 GiB weights), down from 793.6 GB BF16 source (28.5%, −72%) Format compressed tensors mxfp4 pack quantized — per expert, FP4 E2M1, group 32 e8m0 scales, symmetric Base 397B MoE · 512 experts (top 10) + 1 shared · hybrid attention (45 linear attention / 15 full attention) · 60 layers · hidden 4096 · vision encoder · post trained on Qwen 3.5 for agentic coding What is quantized Component Precision Why Routed experts ( mlp.experts. proj ) MXFP4 (4 bit) ~95% of the weights, sparsely activated (top 10/512) — the only place worth the size win Shared expert ( shared expert. ) BF16 dense, runs on every token — quantizing it roughly doubles PPL loss for Note: config.json declares mtp num hidden layers: 1 , but the base release ships no MTP head weights , so multi token speculative decoding is not available with this checkpoint. How…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy