Ornith 1.0 35B — MXFP4 (mixed precision) A 4 bit MXFP4 quantization of Ornith 1.0 35B, produced with qstream . Only the routed MoE experts (~95% of the weights) are quantized to MXFP4; everything quality sensitive — including the always on shared expert — stays BF16 , bit identical to the source. The original model card follows in full below. Size 22.9 GB (21.3 GiB weights), down from 70 GB BF16 source (33%, −67%) Format compressed tensors mxfp4 pack quantized — per expert, FP4 E2M1, group 32 e8m0 scales, symmetric Base 35B A3B MoE · 256 experts (top 8) + 1 shared · hybrid attention (30 linear attention / 10 full attention) · 40 layers · 262 K context · vision encoder · post trained on Qwen 3.5 for agentic coding What is quantized Component Precision Why Routed experts ( mlp.experts. proj.weight ) MXFP4 (4 bit) ~95% of the weights, sparsely activated (top 8/256) — the only place worth the size win Shared expert ( shared expert. proj.weight ) BF16 dense, runs on every token — quantizing it cost ~half the PPL loss for Ornith 1.0 35B Aloha! 🌺 Today, we are releasing Ornith 1.0, a self improving family of open source models for agentic coding. Highlights: State of the Art Coding Agent…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy