SmolLM 1.7B Instruct Model Summary SmolLM is a series of small language models available in three sizes: 135M, 360M, and 1.7B parameters. These models are pre trained on SmolLM Corpus, a curated collection of high quality educational and synthetic data designed for training LLMs. For further details, we refer to our blogpost. To build SmolLM Instruct, we finetuned the base models on publicly available datasets. Changelog Release Description v0.1 Initial release of SmolLM Instruct. We finetune on the permissive subset of the WebInstructSub dataset, combined with StarCoder2 Self OSS Instruct. Then, we perform DPO (Direct Preference Optimization) for one epoch on HelpSteer for the 135M and 1.7B models, and argilla/dpo mix 7k for the 360M model. v0.2 We changed the finetuning mix to datasets more suitable for smol models. We train on a new dataset of 2k simple everyday conversations we generated by llama3.1 70B everyday conversations llama3.1 2k, Magpie Pro 300K Filtered, StarCoder2 Self OSS Instruct, and a small subset of OpenHermes 2.5 v0.2 models are better at staying on topic and responding appropriately to standard prompts, such as greetings and questions about their role as AI as…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy