s batman/Ornith 1.0 35B NVFP4 MTP GGUF MXFP4 + Multi Token Prediction (MTP) heads grafted from Qwen3.6 35B A3B, packaged as a single GGUF for llama.cpp. Designed for NVIDIA Blackwell GPUs (sm 120 / sm 121) including the RTX PRO 6000 and the DGX Spark (GB10) . MXFP4 is dequantized natively by Blackwell tensor cores, and the grafted MTP heads enable draft mtp speculative decoding for ~2× decode throughput versus the body only quant. Original Model Ornith 1.0 35B is a self improving agentic coding model released by the DeepReinforce team, post trained via RL on top of Qwen3.5 35B A3B. It emits a reasoning content block before its final answer and is competitive with Qwen3.6 35B A3B and Gemma 4 31B on Terminal Bench 2.1, SWE bench Verified/Pro, and Claw eval. Architecture: Qwen3.5 MoE ( qwen3 5 moe ), 40 layers, 256 experts, hidden size 2048 Parameters: 35B total / ~3B active Vocabulary: 248,064 tokens (multimodal vocab preserved; vision tower not included in this GGUF) License: MIT (inherited from upstream) Citation: see Citation below Quantization Details This repository ships two files with the same trunk weights but different expert quantizations: File 1 — ornith 1.0 35b MXFP4 MOE…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy