TinyLlama 1.1B Chat v1.0 This repo contains model files for TinyLlama 1.1B Chat v1.0 optimized for nm vllm, a high throughput serving engine for compressed LLMs. This model was quantized with GPTQ and saved in the Marlin format for efficient 4 bit inference. Marlin is a highly optimized inference kernel for 4 bit models. Inference Install nm vllm for fast inference and low memory usage: Run in a Python pipeline for local inference: Quantization For details on how this model was quantized and converted to marlin format, run the quantization/apply gptq save marlin.py script: Slack For further support, and discussions on these models and AI in general, join Neural Magic's Slack Community
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy