Dedicated to building a more intuitive, comprehensive, and efficient LLMs compression toolkit. 📖 Documentation      🤗 Hugging Face      🤖 ModelScope      💬 WeChat 📣Latest News [26/01/13] We have released v0.3. We support the training and deployment of Eagle3 for all scale LLMs/VLMs/Audio models, as detailed in the guidance documentation. And We released Sherry , the hardware efficient 1.25 bit quantization algorithm [Paper Comming soon] [[Code]](https://github.com/Tencent/AngelSlim/tree/sherry/Sherry)🔥🔥🔥 [25/11/05] We have released v0.2. Quantization support for new models, such as GLM 4.6 , Qwen3 VL and Qwen3 Omni , open sources the Eagle3 speculative decoding training framework, and updates the Diffusion model quantization tools. [25/09/30] We have released SpecExit , the reasoning early exit algorithm: [[Paper]](http://arxiv.org/abs/2509.24248) [[Docs]](https://angelslim.readthedocs.io/zh cn/latest/features/speculative decoding/spec exit.html) [[vLLM Code]](https://github.com/vllm project/vllm/pull/27192) [25/09/26] We have released TEQUILA , the ternary quantization algorithm [[Paper]](https://arxiv.org/abs/2509.23809) [[C…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy