SmolTalk Dataset description This is a synthetic dataset designed for supervised finetuning (SFT) of LLMs. It was used to build SmolLM2 Instruct family of models and contains 1M samples. More details in our paper https://arxiv.org/abs/2502.02737 During the development of SmolLM2, we observed that models finetuned on public SFT datasets underperformed compared to other models with proprietary instruction datasets. To address this gap, we created new synthetic datasets that improve instruction following while covering diverse tasks including text editing, rewriting, summarization, and reasoning. Through a series of data ablations at 1.7B scale, we enhanced our SFT mix by incorporating public datasets to strengthen specific capabilities such as mathematics, coding, system prompt following and long context understanding. All the new datasets were generated with distilabel and you can find the generation code here https://github.com/huggingface/smollm/tree/main/text/data/smoltalk. You can load a dataset using Dataset composition The mix consists of: New datasets Smol Magpie Ultra : the core component of our mix, consisting of 400K samples generated using the Magpie pipeline with /Llama…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy