MegaBeam Mistral 7B 300k Model MegaBeam Mistral 7B 300k is a fine tuned Mistral 7B Instruct v0.2 language model that supports input contexts up to 320k tokens. MegaBeam Mistral 7B 300k can be deployed on a single AWS g5.48xlarge instance using serving frameworks such as vLLM, Sagemaker DJL endpoint, and others. Similarities and differences beween MegaBeam Mistral 7B 300k and Mistral 7B Instruct v0.2 are summarized below: Model Max context length rope theta prompt template : : : Mistral 7B Instruct v0.2 32K 1e6 instruction format MegaBeam Mistral 7B 300k 320K 25e6 AS ABOVE Evaluations InfiniteBench: Extending Long Context Evaluation Beyond 100K Tokens InfiniteBench is a cutting edge benchmark tailored for evaluating the capabilities of language models to process, understand, and reason over super long contexts (100k+ tokens) . We therefore evaluated MegaBeam Mistral 7B 300k, Mistral 7B Instruct v0.2, Llama 3 8B Instruct 262k, and Llama3 70B 1M on InfiniteBench. The InfiniteBench authors also evaluated SOTA proprietary and open source LLMs on InfiniteBench. We thus combined both results in the table below. Task Name MegaBeam Mistral 7B 300k Mistral 7B Instruct v0.2 Llama 3 8B Instruc…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy