Model Card for Ministral 8B Instruct 2410 We introduce two new state of the art models for local intelligence, on device computing, and at the edge use cases. We call them les Ministraux: Ministral 3B and Ministral 8B. The Ministral 8B Instruct 2410 Language Model is an instruct fine tuned model significantly outperforming existing models of similar size, released under the Mistral Research License. If you are interested in using Ministral 3B or Ministral 8B commercially, outperforming Mistral 7B, reach out to us. For more details about les Ministraux please refer to our release blog post. Ministral 8B Key features Released under the Mistral Research License , reach out to us for a commercial license Trained with a 128k context window with interleaved sliding window attention Trained on a large proportion of multilingual and code data Supports function calling Vocabulary size of 131k , using the V3 Tekken tokenizer Basic Instruct Template (V3 Tekken) For more information about the tokenizer please refer to mistral common Ministral 8B Architecture Feature Value : : : : Architecture Dense Transformer Parameters 8,019,808,256 Layers 36 Heads 32 Dim 4096 KV Heads (GQA) 8 Hidden Dim 122…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy