Qwen3.5 9B MTP GGUF Qwen3.5 9B from Alibaba's Qwen team is a 9B parameter dense multimodal language model featuring a hybrid Gated DeltaNet + Gated Attention architecture with 262K native context window (extensible to 1M+ tokens via RoPE scaling), 248K vocabulary supporting 201 languages, and early‑fusion training for unified text, image, and video understanding. It achieves SOTA performance across modalities with 89.2% OCRBench, 84.5% VideoMME, 78.9% MathVision, and 70.1% MMMU Pro, while delivering production level agentic capabilities including 66.1% BFCL V4 and 79.1% TAU2 Bench for native tool calling, plus toggleable thinking mode for step by step reasoning on complex tasks. Apache 2.0 licensed and optimized for vLLM/SGLang/llama.cpp/Ollama deployment (~18GB VRAM BF16, ~5GB 4 bit), the instruction tuned variant excels at repository level coding, frontend development, document/PDF parsing, visual question answering, and multilingual chatbots as a scalable foundation for edge to server multimodal agents. [!NOTE] Multi Token Prediction (MTP) GGUF is a specialized GGUF model file format extension that integrates speculative decoding directly into the model weights to significantly…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy