PLaMo 2 1B Model Description PLaMo 2 1B is a 1B model pre trained on English and Japanese datasets, developed by Preferred Elements, Inc. PLaMo 2 models adapt the hybrid architecture like Samba rather than the Transformer architecture. Samba integrates Mamba, a selective State Space Model (SSM), with sliding window attention, combining their strengths for improved efficiency and performance. The major differences between Samba and PLaMo 2 are 1) adding normalization layers to improve training stability, and 2) using Mamba2 kernel for computational efficiency. PLaMo 2 1B is released under Apache License version 2.0. NOTE : This model has NOT been instruction tuned for chat dialog or other downstream tasks. Usage Requirements Use a pipeline as a high level helper Load model directly Model Details Model size: 1B Trained tokens: 4T tokens Developed by: Preferred Elements, Inc. Model type: Causal decoder only Language(s): English, Japanese License: Apache License version 2.0 Training Dataset We trained PLaMo 2 1B in two phases, phase 1 with 3.5T tokens and phase 2 with 0.5T tokens. The percentage of datasets in each phase is shown in the following table. 3.5T (phase 1) 0.5T (phase 2) To…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy