Warning, all models only work with ik llama.cpp Quants of Ornith 1.0, a fine tune built on Qwen 3.5 397B A17B. Comes with mmproj for vision, but isn't shipped with MTP. You can use DFLASH with it, a novel diffusion based MTP like, to speed up TG comes in a variety of quants, you can download the one that works best for your model size. DFLASH paper: https://arxiv.org/abs/2602.06036 Thanks to: https://huggingface.co/z lab/Qwen3.5 397B A17B DFlash https://huggingface.co/modal labs/Qwen3.5 397B A17B DFlash https://huggingface.co/lmsys/Qwen3.5 397B A17B DFlash Load DFLASH with: All quants target 16/24/32GB GPUs, with varying amounts of RAM depending on the quant. Specific quant details (memory footprint with mmproj, without MTP/DFLASH): IQ4 K for 256GB RAM + 24GB VRAM Will eat 20180MB of VRAM and 198GB of RAM with standard config: Details: IQ4 KSS for 256GB RAM + 24GB VRAM Will eat 18826MB of VRAM and 191GB of RAM with standard config: Details: IQ3 KS for 192GB RAM + 24GB VRAM Will eat 17600MB of VRAM and 137GB of RAM with standard config: Details: IQ2 KS for 128GB RAM + 16GB VRAM Will eat 13988MB of VRAM and 92.4GB of RAM with standard config: Details: Every additional 65536 tokens of…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy