News Our first data centric LLM competition begins! Please visit the competition's official websites, FT Data Ranker (1B Track, 7B Track), for more information. Introduction This is a reference LLM from Data Juicer. The model architecture is LLaMA 1.3B and we adopt the OpenLLaMA implementation. The model is pre trained on 150B tokens of Data Juicer's refined RedPajama and Pile. It achieves an average score of 34.21 over 16 HELM tasks, beating Falcon 1.3B (trained on 350B tokens from RefinedWeb), Pythia 1.4B (trained on 300B tokens from original Pile) and Open LLaMA 1.3B (trained on 150B tokens from original RedPajama and Pile). For more details, please refer to our paper.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy