SmolLM3 GGUF Original model: https://huggingface.co/HuggingFaceTB/SmolLM3 3B [!IMPORTANT] To enable thinking, you need to specify jinja Example usage with llama.cpp: Table of Contents 1. Model Summary 2. Evaluation 3. Training 4. Limitations 5. License Model Summary SmolLM3 is a 3B parameter language model designed to push the boundaries of small models. It supports 6 languages, advanced reasoning and long context. SmolLM3 is a fully open model that offers strong performance at the 3B–4B scale. The model is a decoder only transformer using GQA and NoRope, it was pretrained on 11.2T tokens with a staged curriculum of web, code, math and reasoning data. Post training included midtraining on 140B reasoning tokens followed by supervised fine tuning and alignment via Anchored Preference Optimization (APO). Key features Instruct model optimized for hybrid reasoning Fully open model : open weights + full training details including public data mixture and training configs Long context: Trained on 64k context and suppots up to 128k tokens using YARN extrapolation Multilingual : 6 natively supported (English, French, Spanish, German, Italian, and Portuguese) For more details refer to our blo…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy