See our collection for versions of Llama 4 including 4 bit & 16 bit formats. 🦙 Run Unsloth Dynamic Llama 4 GGUF! Read our Guide to see how to Fine tune & Run Llama 4 correctly. MoE Bits Type Disk Size HF Link Accuracy : : : : : 1.78bit IQ1\ S 122GB Link Ok 1.93bit IQ1\ M 128GB Link Fair 2.42 bit IQ2\ XXS 140GB Link Better 2.71 bit Q2\ K\ XL 151B Link Suggested 3.5 bit Q3\ K\ XL 193GB Link Great 4.5 bit Q4\ K\ XL 243GB Link Best Currently text only is supported. Chat template/prompt format: 🦙 Fine tune Meta's Llama 4 with Unsloth! Fine tune Llama 4 Scout on a single H100 80GB GPU using Unsloth! Read our Blog about Llama 4 support: unsloth.ai/blog/llama4 View the rest of our notebooks in our docs here. Export your fine tuned model to GGUF, Ollama, llama.cpp, vLLM or 🤗HF. Unsloth supports Free Notebooks Performance Memory use GRPO with Llama 3.1 (8B) ▶️ Start on Colab GRPO.ipynb) 2x faster 80% less Llama 3.2 (3B) ▶️ Start on Colab Conversational.ipynb) 2.4x faster 58% less Llama 3.2 (11B vision) ▶️ Start on Colab Vision.ipynb) 2x faster 60% less Qwen2.5 (7B) ▶️ Start on Colab Alpaca.ipynb) 2x faster 60% less Phi 4 (14B) ▶️ Start on Colab 2x faster 50% less Mistral (7B) ▶️ Start on…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy