MistralLite Model MistralLite is a fine tuned Mistral 7B v0.1 language model, with enhanced capabilities of processing long context (up to 32K tokens). By utilizing an adapted Rotary Embedding and sliding window during fine tuning, MistralLite is able to perform significantly better on several long context retrieve and answering tasks , while keeping the simple model structure of the original model. MistralLite is useful for applications such as long context line and topic retrieval, summarization, question answering, and etc. MistralLite can be deployed on a single AWS g5.2x instance with Sagemaker Huggingface Text Generation Inference (TGI) endpoint, making it suitable for applications that require high performance in resource constrained environments. You can also serve the MistralLite model directly using TGI docker containers. Also, MistralLite supports other ways of serving like vLLM, and you can use MistralLite in Python by using the HuggingFace transformers and FlashAttention 2 library. MistralLite is similar to Mistral 7B Instruct v0.1, and their similarities and differences are summarized below: Model Fine tuned on long contexts Max context length RotaryEmbedding adaptati…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy