See our collection for versions of Llama 3.1 including 4 bit & 16 bit formats. Unsloth Dynamic v2.0 achieves superior accuracy & outperforms other leading quant methods. 🦙 Run Unsloth Llama 3.1 GGUF! Read our Blog about Llama 3.1 fine tuning support: unsloth.ai/blog/llama4 View the rest of our fine tuning notebooks in our docs here. Export your fine tuned model to GGUF, Ollama, llama.cpp, vLLM or 🤗HF. Unsloth supports Free Notebooks Performance Memory use GRPO with Llama 3.1 (8B) ▶️ Start on Colab GRPO.ipynb) 2x faster 80% less Llama 3.2 (3B) ▶️ Start on Colab Conversational.ipynb) 2.4x faster 58% less Llama 3.2 (11B vision) ▶️ Start on Colab Vision.ipynb) 2x faster 60% less Qwen2.5 (7B) ▶️ Start on Colab Alpaca.ipynb) 2x faster 60% less Phi 4 (14B) ▶️ Start on Colab 2x faster 50% less Mistral (7B) ▶️ Start on Colab Conversational.ipynb) 2.2x faster 62% less Llama 3.1 Model Information The Meta Llama 3.1 collection of multilingual large language models (LLMs) is a collection of pretrained and instruction tuned generative models in 8B, 70B and 405B sizes (text in/text out). The Llama 3.1 instruction tuned text only models (8B, 70B, 405B) are optimized for multilingual dialogue use…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy