InternVL2 2B AWQ [\[๐ GitHub\]](https://github.com/OpenGVLab/InternVL) [\[๐ InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[๐ InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[๐ Mini InternVL\]](https://arxiv.org/abs/2410.16261) [\[๐ InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[๐ Blog\]](https://internvl.github.io/blog/) [\[๐จ๏ธ Chat Demo\]](https://internvl.opengvlab.com/) [\[๐ค HF Demo\]](https://huggingface.co/spaces/OpenGVLab/InternVL) [\[๐ Quick Start\]]( quick start) [\[๐ Documents\]](https://internvl.readthedocs.io/en/latest/) Introduction INT4 Weight only Quantization and Deployment (W4A16) LMDeploy adopts AWQ algorithm for 4bit weight only quantization. By developed the high performance cuda kernel, the 4bit quantized model inference achieves up to 2.4x faster than FP16. LMDeploy supports the following NVIDIA GPU for W4A16 inference: Turing(sm75): 20 series, T4 Ampere(sm80,sm86): 30 series, A10, A16, A30, A100 Ada Lovelace(sm90): 40 series Before proceeding with the quantization and inference, please ensure that lmdeploy is installed. This article comprises the following sections: Inference Service Inference Trying the followiโฆ
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy