Ornith 1.0 35B — PrismaAURA 4.75 bit (mixed NVFP4/FP8/BF16) + MTP Mixed precision quantization of deepreinforce ai/Ornith 1.0 35B via PrismaAURA — per Linear bit allocation selected on real end to end KL, shipped as a stock compressed tensors checkpoint that vanilla vLLM serves with no forked runtime and no custom kernels . Includes a Multi Token Prediction (MTP) head for speculative decoding. Metric Value Effective rate (body) 4.75 bits / quantizable parameter Size ~23 GB (from ~70 GB BF16) Served confident KL vs BF16 0.0143 (top 1 agreement 98.6%) MTP acceptance 91.3% @ pos 0, 77.3% overall (~2.32 accepted / 3 drafted) Serving vanilla vLLM, CUTLASS native NVFP4 W4A4 on Blackwell MTP note The base Ornith checkpoint ships no MTP weights (the fine tune stripped them). This artifact grafts the base Qwen3.5 35B A3B MTP head (BF16), the same head every community MTP Ornith variant uses. Because it predicts the base distribution rather than Ornith's RL tuned one, acceptance is a bit lower than a native head would give — but still strong (91% at position 0). Serving With speculative decoding (MTP): Without MTP (e.g. for clean perplexity measurement — spec decode poisons logprob metrics),…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy