MTP GGUF QwenPaw Flash 9B heretic MTP English π δΈζζζ‘£ QwenPaw Flash 9B heretic non MTP Version: QwenPaw Flash 9B heretic GGUF π BenchLocal Total: 4035/5000 (80.7%) β MTP Speculative Decoding Injected Uncensored Β· Abliterated Β· Agent Optimized Β· 1.7 4.1Γ Speedup Uncensored version of QwenPaw Flash 9B , processed with Heretic v1.3.0 abliteration, with MTP (Multi Token Prediction) head weights injected from the original Qwen3.5 9B base model. By reconstructing the MTP speculative decoding head β which was stripped during the QwenPaw fine tuning process β this model achieves up to 4.1Γ inference speedup on real agent benchmarks while maintaining or improving accuracy. π π BenchLocal Benchmarks (With MTP) Test Environment : NVIDIA RTX 5070 Ti (16GB) Β· llama.cpp (turboquant build, spec type draft mtp ) Β· Q6 K quant Framework : BenchLocal β local model agent evaluation suite Methodology : Each scenario run once , no retries, no second attempts Benchmark Score Accuracy Results Time vs No MTP ToolCall 15 π οΈ 1500/1500 100% 15β 0β οΈ 0β 0.65min 1.4Γ faster HermesAgent 20 π€ 1505/2000 75.3% 12β 1β οΈ 7β 5.3min 1.17Γ faster BugFind 15 π 1030/1500 68.7% 9β 2β οΈ 4β 1.8min 4.1Γ faster Total 4035/5β¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy