Ornith 1.0 35B AEON Ultimate Uncensored NVFP4 NVFP4 (4 bit) build of the uncensored … AEON Ultimate Uncensored BF16 — itself an abliterated build of deepreinforce ai/Ornith 1.0 35B , DeepReinforce's SOTA agentic coding MoE. 23.7 GB (64 % smaller than BF16), near lossless , refusals removed. What this quant is (and isn't) A surgically scoped MLP only, weight only NVFP4 quant: Component Precision Why MoE experts (all 256) + shared expert MLP NVFP4 (W4A16) ~90 % of the weights; where the size lives Full attention (q/k/v/o) BF16 preserved — FP4 attention costs reasoning GatedDeltaNet / SSM ( linear attn. ) BF16 recurrent path must stay full precision Vision tower ( model.visual. ) BF16 multimodal kept intact Router gates, embeddings, lm head, norms BF16 routing/IO precision sensitive Weight only (W4A16): weights are NVFP4, activations stay BF16 . This is the key choice — the well known NVFP4 reasoning penalty comes from FP4 activations (W4A4), not the weights. Weight only sidesteps it and, on Blackwell, dequantizes to BF16 in register (also avoiding the sm 120/121 grouped FP4 GEMM failure mode). Validation (measured on the served NVFP4 model, vs the BF16 parent) Metric BF16 parent This…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy