🍮 The WHOLE FLAN Collection! 🍮 Overview This repository includes the full dataset from the FLAN Collection, totalling ~300GB as parquets. Generated using the official seqio templating from the Google FLAN Collection GitHub repo. The data is subject to all the same licensing of the component datasets. To keep up with our continued work on OpenOrca and other exciting research, find our Discord here: https://AlignmentLab.ai Motivation This work was done as part of the requirements for the OpenOrca project. There was not a large enough subset of FLAN Collection generated publicly to subsample from to complete the work. So, we opted to process the entire collection ourselves. Generating this requires an understanding of seqio and a Linux server with 512GB of CPU ram, as well as fast drives and custom limits for many parameters beyond what is default on Linux server distributions (e.g., requiring up to 45,000 threads running at once). It takes downloading over 400GB of datasets, working around tfds bugs, and then processing the datasets over the course of several days. We provide this repo as a resource to other ML researchers, as it saves these time consuming and laborious steps to ge…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy