Dedicated to building a more intuitive, comprehensive, and efficient LLMs compression toolkit. 📖 Documentation      🤗 Hugging Face      🤖 ModelScope      💬 WeChat Table of Contents Latest Updates Key Features Supported Models How to Use Install AngelSlim Quick Start deployment & Evaluation Benchmark License Citation Technical Discussion 📣Latest Updates [25/11/05] We have released v0.2. Quantization support for new models, such as GLM 4.6 , Qwen3 VL and Qwen3 Omni , open sources the Eagle3 speculative decoding training framework, and updates the Diffusion model quantization tools. [25/09/30] We have released SpecExit , the reasoning early exit algorithm: [[Paper]](http://arxiv.org/abs/2509.24248) [[Docs]](https://angelslim.readthedocs.io/zh cn/latest/features/speculative decoding/spec exit.html) [[vLLM Code]](https://github.com/vllm project/vllm/pull/27192)🔥🔥🔥 [25/09/26] We have released TEQUILA , the ternary quantization algorithm [[Paper]](https://arxiv.org/abs/2509.23809) [[Code]](https://github.com/Tencent/AngelSlim/tree/tequila/TernaryQuant)🔥🔥🔥 [25/09/24] We now support the PTQ quantification of NVFP4 for the Qwen3 seri…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy