Cosmopedia v0.1 Image generated by DALL E, the prompt was generated by Mixtral 8x7B Instruct v0.1 Note: Cosmopedia v0.2 is available at smollm corpus Cosmopedia is a dataset of synthetic textbooks, blogposts, stories, posts and WikiHow articles generated by Mixtral 8x7B Instruct v0.1.The dataset contains over 30 million files and 25 billion tokens , making it the largest open synthetic dataset to date. It covers a variety of topics; we tried to map world knowledge present in Web datasets like RefinedWeb and RedPajama, and generate synthetic content that covers them. This is the v0.1 of Cosmopedia, with ample room for improvement and topics to be more comprehensively covered. We hope this dataset will help the community's research efforts in the increasingly intriguing domain of synthetic data. You can find a clickable map by Nomic at https://atlas.nomic.ai/map/cosmopedia. This work is inspired by the great work of Phi1.5. You can find more details about the dataset in our blog post : https://huggingface.co/blog/cosmopedia TL;DR This is a synthetic dataset of 30M samples generated by Mixtral 8x7B Instruct v0.1. It contains 8 splits depending on the source of the seed samples we use…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy