Instruction Finetuning Dataset Collection (Alpaca CoT) This repository will continuously collect various instruction tuning datasets. And we standardize different datasets into the same format, which can be directly loaded by the code of Alpaca model. We also have conducted empirical study on various instruction tuning datasets based on the Alpaca model, as shown in https://github.com/PhoebusSi/alpaca CoT. If you think this dataset collection is helpful to you, please like this dataset and star our github project! You are in a warm welcome to provide us with any non collected instruction tuning datasets (or their sources). We will uniformly format them, train Alpaca model with these datasets and open source the model checkpoints. Contribute Welcome to join us and become a contributor to this project! If you want to share some datasets, adjust the data in the following format: Folder should be like this: Create a new pull request in Community and publish your branch when you are ready. We will merge it as soon as we can. Data Usage and Resources Data Format All data in this folder is formatted into the same templates, where each sample is as follows: alpaca alpaca data.json This dat…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy