Table of Contents Dataset Summary Dataset Attribution Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Dataset Use Use Cases Usage Caveats Getting Started 🐋 The OpenOrca Dataset! 🐋 We are thrilled to announce the release of the OpenOrca dataset! This rich collection of augmented FLAN data aligns, as best as possible, with the distributions outlined in the Orca paper. It has been instrumental in generating high performing model checkpoints and serves as a valuable resource for all NLP researchers and developers! Official Models Mistral 7B OpenOrca Our latest model, the first 7B to score better overall than all previous models below 30B. 98% of Llama2 70b chat's performance, in a completely open 7B! OpenOrca Platypus2 13B Our third model, the first 13B model to score higher than LLaMA1 65B on the HuggingFace Leaderboard! Released in partnership with Platypus. LlongOrca 7B & 13B Our first 7B release, trained on top of LLongMA2 to achieve 16,000 tokens context. 1 long context 7B model at release time, with 99% of the overall 1 model's performance. LlongOrca 13B 16k, trained on top of LL…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy