Nex N2 mini GGUF GGUF quantizations of nex agi/Nex N2 mini for use with llama.cpp. These are Unsloth style UD (dynamic) quants : per tensor quantization types tuned with an importance matrix, using the same recipe family as Unsloth’s Qwen3.6 35B A3B MoE GGUF releases. Model at a glance Architecture qwen35moe (Qwen3.5 / 3.6 MoE family) Trunk layers 40 Experts 256 total, 8 active per token Context (train) 262144 tokens Vocab 248320 Vision Supported via mmproj BF16.gguf (optional) MTP draft head Not included in this release (see note below) Files File Size When to use : Nex N2 mini UD Q3 K XL.gguf ~17 GB Smallest; more VRAM friendly Nex N2 mini UD Q4 K M.gguf ~22 GB Good default balance Nex N2 mini UD Q4 K XL.gguf ~22 GB Recommended quality / size sweet spot Nex N2 mini UD Q5 K XL.gguf ~27 GB Higher quality Nex N2 mini UD Q6 K XL.gguf ~32 GB Highest quality in this set mmproj BF16.gguf ~0.9 GB Image / vision input (optional) imatrix unsloth.gguf file ~0.2 GB Importance matrix used during quantization (reference only) All .gguf model files are at the repo root (flat layout). Quick start Chat server (recommended) Open http://127.0.0.1:8080 in your browser for the built in chat UI. CLI V…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy