Ornith 1.0 35B: 1M Context + MTP + Vision 181717?logo=github) Mirrors: Hugging Face ModelScope (full quant ladder always available on ModelScope) DeepReinforce's Ornith 1.0 35B (35B MoE, 3B active, Qwen3.5 family) with YaRN RoPE scaling baked into the GGUF metadata for a 1,048,576 token context window , needle certified through the full range, now shipping MTP speculative decoding and a vision tower. Weights are bit identical to the source builds; only rope metadata differs, so llama.cpp and Ollama apply the extension with no extra flags. Capability Status Evidence in this repo 1M context Certified 50/50 needles through 1M , f16 KV (heatmap + results.jsonl) MTP speculative decoding Two vetted builds Q6 K: 267.9 to 382.9 tok/s ( +43% ). APEX Compact: +14%, needle perfect 70/70 through 1M (524K on 5090, 786K/1M on M3 Max) Vision Verified mmproj tower reads image text and identifies objects correctly Files: MTP first File Size Pick it when ornith 1.0 35b 1M MTP Q6 K.gguf 29.2 GB You have the memory (64 GB Mac, RTX Pro 6000). Fastest: +43% decode, Q6 quality ornith 1.0 35b 1M MTP APEX Compact.gguf 17.0 GB You want 1M class context on a 32 GB card. SC117 APEX imatrix quant, needle perfe…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy