GLM 5.2 — colibrì int4 container (~370 GB) This is the EXACT SAME THING as https://huggingface.co/jlnsrk/GLM 5.2 colibri int4, BUT with int8 MTP heads, which are needed for speculative decoding—and with that, an overall major inference speedboost. The original int4 MTP heads have low acceptance rate, and are essentially useless. ⚠️ This is NOT a GGUF / AWQ / GPTQ / MLX model. It only works with the colibrì engine. Usage Requirements: Linux (or WSL2), gcc + OpenMP, AVX2, ≥16 GB RAM, ~400 GB free NVMe. Provenance & license Converted from zai org/GLM 5.2 FP8 (MIT). This derivative is likewise MIT. Conversion performed with colibrì's official converter, unmodified. Cloned & modded from jlnsrk/GLM 5.2 colibri int4
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy