SmolTalk2 Dataset description This dataset contains three subsets (Mid, SFT, Preference) that correspond to the three phases of Post Training for SmolLM3 3B. You can find more details in our blog post about how we used the data in each of the stages SmolLM3. The specific weight of each subset is available in the training recipe in SmolLM's repository. You can load a dataset using from datasets import load dataset To load the train split of a specific subset… See the full description on the dataset page: https://huggingface.co/datasets/HuggingFaceTB/smoltalk2.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy