🐉 qwen27B Agent R2 abliterated preview 27B Agent Model — Abliterated · MTP · Tool Calling · Speculative Decoding Preview release — Built from Fable MTP + agent LoRA fusion. Features Multi Token Prediction (MTP) for speculative decoding (up to 2× faster generation), abliterated (no guardrails), and tool calling support. ✨ Key Features Capability Description ⚡ MTP Speculative Decoding Draft 2 tokens at a time — up to +85% decode TPS on single GPU 🔧 Tool Calling Hermes/Qwen function calling format via llama.cpp tools all 🔓 Abliterated Unrestricted — all refusal mechanisms removed 🧠 Reasoning Fable style reasoning with step by step CoT 🌏 Thai + English Native bilingual support 💻 Code Python, shell, system tasks 🚀 Usage llama.cpp (Recommended) Parameter Purpose cache type k bf16 / cache type v bf16 BF16 KV cache for quality flash attn on Flash attention for speed tools all Enable tool/function calling spec type draft mtp MTP speculative decoding (draft 2 tokens) spec draft n max 2 Max 2 draft tokens per step cont batching Continuous batching for multi turn jinja Use Jinja2 chat template from GGUF Python (Transformers) 📦 Downloads File Size Quant Description : : : : qwen27B Agent…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy