Ornith 1.0 35B — APEX GGUF APEX (Adaptive Precision for EXpert Models) quantizations of Ornith 1.0 35B, an open source coding MoE model by DeepReinforce (MIT license, based on Qwen 3.5 architecture). These quants were produced using the apex quant toolchain. APEX is a MoE aware mixed precision quantization strategy that classifies tensors by role (routed expert, shared expert, attention) and applies a layer wise precision gradient — edge layers get higher precision, middle layers more aggressive compression. Files Each profile comes in two variants: Base — the quantized model standalone MTP — includes the bundled MTP (multi token prediction) head, quantized to Q8 0 (near lossless), for self speculative decoding via spec type draft mtp . Requires a recent llama.cpp build with MTP support. I variants were calibrated with a diverse importance matrix (chat, code, reasoning, tool calling, multilingual) for improved downstream accuracy. File Profile Size Best For ornith 1.0 35b APEX I Mini.gguf I Mini 14 GB Smallest viable, fastest inference ornith 1.0 35b APEX I Mini MTP.gguf I Mini + MTP 14 GB Smallest viable + self spec ornith 1.0 35b APEX Compact.gguf Compact 17 GB Consumer GPUs, gen…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy