EXAONE 3.5 32B Instruct AWQ Introduction We introduce EXAONE 3.5, a collection of instruction tuned bilingual (English and Korean) generative models ranging from 2.4B to 32B parameters, developed and released by LG AI Research. EXAONE 3.5 language models include: 1) 2.4B model optimized for deployment on small or resource constrained devices, 2) 7.8B model matching the size of its predecessor but offering improved performance, and 3) 32B model delivering powerful performance. All models support long context processing of up to 32K tokens. Each model demonstrates state of the art performance in real world use cases and long context understanding, while remaining competitive in general domains compared to recently released models of similar sizes. For more details, please refer to our technical report, blog and GitHub. This repository contains the AWQ quantized weights of the instruction tuned 32B language model with the following features: Number of Parameters (without embeddings): 30.95B Number of Layers: 64 Number of Attention Heads: GQA with 40 Q heads and 8 KV heads Vocab Size: 102,400 Context Length: 32,768 tokens Quantization: AWQ with 4 bit group wise weight only quantization…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy