s batman/Ornith 1.0 9B NVFP4 MTP GGUF Three honest quantizations of deepreinforce ai/Ornith 1.0 9B with Multi Token Prediction (MTP) heads grafted from Qwen3.5 9B MTP , packaged as three GGUF variants for llama.cpp. Designed for NVIDIA Blackwell GPUs (sm 120 / sm 121) including the RTX PRO 6000 and the DGX Spark (GB10) . NVFP4 and MXFP4 are dequantized natively by Blackwell tensor cores; the grafted MTP heads enable draft mtp speculative decoding for significant decode throughput uplift. Naming note (read first): the repo is named NVFP4 MTP GGUF because that was the original (incorrect) name. The repo now hosts three quant files of three different formats — see the Provided Files table. Pick the file whose name matches its actual tensor layout. Original Model Ornith 1.0 9B is a self improving agentic coding model released by the DeepReinforce team, post trained via RL on top of Qwen3.5 9B. It is the dense sibling of Ornith 1.0 35B (which is MoE). Both share the Qwen3.5 hybrid attention trunk (linear SSM + full attention every 4th layer) and the same reasoning content block before answer output format. Architecture: Qwen3.5 dense ( qwen3 5 text ), 32 trunk layers + 1 embedded MTP la…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy