EXAONE 3.5 7.8B Instruct Introduction We introduce EXAONE 3.5, a collection of instruction tuned bilingual (English and Korean) generative models ranging from 2.4B to 32B parameters, developed and released by LG AI Research. EXAONE 3.5 language models include: 1) 2.4B model optimized for deployment on small or resource constrained devices, 2) 7.8B model matching the size of its predecessor but offering improved performance, and 3) 32B model delivering powerful performance. All models support long context processing of up to 32K tokens. Each model demonstrates state of the art performance in real world use cases and long context understanding, while remaining competitive in general domains compared to recently released models of similar sizes. For more details, please refer to our technical report, blog and GitHub. This repository contains the instruction tuned 7.8B language model with the following features: Number of Parameters (without embeddings): 6.98B Number of Layers: 32 Number of Attention Heads: GQA with 32 Q heads and 8 KV heads Vocab Size: 102,400 Context Length: 32,768 tokens Quickstart We recommend to use transformers v4.43 or later. Here is the code snippet to run conv…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy