MegaBeam Mistral 7B 512k Model This model, presented in Scaling Context, Not Parameters: Training a Compact 7B Language Model for Efficient Long Context Processing, is a Long Context LLM that supports 524,288 tokens in its context. MegaBeam Mistral 7B 512k was trained on Mistral 7B Instruct v0.2, and can be deployed using various serving frameworks like vLLM and Amazon SageMaker's DJL endpoint. Please refer to our GitRepo for deployment and inference examples. New update! Watch our talk on MegaBeam at NeurIPS 2024 Evaluations We evaluated MegaBeam Mistral 7B 512k on three long context benchmarks. For each benchmark, we deployed the MegaBeam Mistral 7B 512k model with vLLM (v0.5.1) on an EC2 instance and obtained LLM responses through the OpenAI API provided by vLLM. 1. Needle In A Haystack Pressure Testing LLMs The Arize ai NIAH varies the target random number and introduces a random city for each question, requiring the LLM to extract the random number from various selected context locations. MegaBeam Mistral 7B 512k scored 100% on this NIAH benchmark as shown in this plot. 2. RULER: What’s the Real Context Size of Your Long Context Language Models? The RULER benchmark evaluates l…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy