Llama 3.1 Swallow Built with Llama Llama 3.1 Swallow is a series of large language models (8B, 70B) that were built by continual pre training on the Meta Llama 3.1 models. Llama 3.1 Swallow enhanced the Japanese language capabilities of the original Llama 3.1 while retaining the English language capabilities. We use approximately 200 billion tokens that were sampled from a large Japanese web corpus (Swallow Corpus Version 2), Japanese and English Wikipedia articles, and mathematical and coding contents, etc (see the Training Datasets section of the base model) for continual pre training. The instruction tuned models (Instruct) were built by supervised fine tuning (SFT) on the synthetic data specially built for Japanese. See the Swallow Model Index section to find other model variants. Note : Llama 3.1 Swallow 8B Instruct v0.5 model was continually pre trained from the meta llama/Llama 3.1 8B Instruct and then instruction tuned with our instruction datasets. Release History June 25, 2025 : Released Llama 3.1 Swallow 8B Instruct v0.5 and Llama 3.1 Swallow 8B v0.5. March 10, 2025 : Released Llama 3.3 Swallow 70B Instruct v0.4 and Llama 3.3 Swallow 70B v0.4. December 30, 2024 : Release…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy