See our collection for versions of Llama 4 including 4 bit & 16 bit formats. Unsloth Dynamic v2.0 achieves superior accuracy & outperforms other leading quant methods. 🦙 Run Unsloth Dynamic Llama 4 GGUF! Read our Guide to see how to Fine tune & Run Llama 4 correctly. MoE Bits Type Disk Size HF Link Accuracy : : : : : 1.78bit IQ1\ S 33.8GB Link Ok 1.93bit IQ1\ M 35.4GB Link Fair 2.42 bit IQ2\ XXS 38.6GB Link Better 2.71 bit Q2\ K\ XL 42.2GB Link Suggested 3.5 bit Q3\ K\ XL 52.9GB Link Great 4.5 bit Q4\ K\ XL 65.6GB Link Best Currently text only is supported. Chat template/prompt format: 🦙 Fine tune Meta's Llama 4 with Unsloth! Fine tune Llama 4 Scout on a single H100 80GB GPU using Unsloth! Read our Blog about Llama 4 support: unsloth.ai/blog/llama4 View the rest of our notebooks in our docs here. Export your fine tuned model to GGUF, Ollama, llama.cpp, vLLM or 🤗HF. Unsloth supports Free Notebooks Performance Memory use GRPO with Llama 3.1 (8B) ▶️ Start on Colab GRPO.ipynb) 2x faster 80% less Llama 3.2 (3B) ▶️ Start on Colab Conversational.ipynb) 2.4x faster 58% less Llama 3.2 (11B vision) ▶️ Start on Colab Vision.ipynb) 2x faster 60% less Qwen2.5 (7B) ▶️ Start on Colab Alpaca.ip…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy