princeton nlp/Llama 3 8B ProLong 64k Instruct [Paper] [HF Collection] [Code] ProLong ( Pr incet o n long context language models) is a family of long context models that are continued trained and supervised fine tuned from Llama 3 8B, with a maximum context window of 512K tokens. Our main ProLong model is one of the best performing long context models at the 10B scale (evaluated by HELMET). To train this strong long context model, we conduct thorough ablations on the long context pre training data, SFT data, and numerous other design choices. We demonstrate our findings in our paper, How to Train Long Context Language Models (Effectively). Authors: Tianyu Gao\ , Alexander Wettig\ , Howard Yen, Danqi Chen ( equal contribution) Contact: {tianyug, awettig}@princeton.edu The ProLong Models princeton nlp/Llama 3 8B ProLong 64k Base princeton nlp/Llama 3 8B ProLong 64k Instruct ← you are here! princeton nlp/Llama 3 8B ProLong 512k Base ⭐ princeton nlp/Llama 3 8B ProLong 512k Instruct Model card Here are some quick facts about our main ProLong model: princeton nlp/Llama 3 8B ProLong 512k Instruct. Base model: meta llama/Meta Llama 3 8B Instruct Long context continued training: 20B tokens…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy