Qwen3.5 9b Sushi Coder RL GGUF Lineage Base model lineage: bigatuna/Qwen3.5 9b Sushi Coder RL model: bigatuna/Qwen3.5 9b Sushi Coder RL RL pipeline: NousResearch/atropos Training The upstream SFT model was trained with Unsloth on: nohurry/Opus 4.6 Reasoning 3000x filtered open r1/codeforces cots The RL stage was then run for coding with NousResearch/hermes agent using NousResearch/atropos. During that run, vLLM was patched with vllm project/vllm PR 36395, fix(lora): add bounds checking for TP configurations , to address the LoRA tensor parallel bounds issue. Files Qwen3.5 9b Sushi Coder RL.Q4 K M.gguf Qwen3.5 9b Sushi Coder RL.Q8 0.gguf Qwen3.5 9b Sushi Coder RL.BF16 mmproj.gguf Usage Note This is a multimodal Qwen 3.5 export. Use the text GGUF together with the BF16 mmproj file. Quick Start Example download commands with the Hugging Face CLI: Alternative quant: Metadata License: Apache 2.0 Architecture: Qwen 3.5 Format: GGUF Tags: llama.cpp , qwen3 5 , multimodal , code , rl , conversational LiveCodeBench Evaluation The benchmark results below were produced from matched local BF16 vLLM endpoints so the RL model and the base model were evaluated with the same serving method, the sa…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy