Model Card for Zamba 7B Zamba 7B v1 is a hybrid model between Mamba, a state space model, and transformers. It uses a mamba backbone with a shared transformer layer every 6 blocks. Zamba was trained using next token prediction. It uses the Mistral v0.1 tokenizer. We came to this architecture after a series of ablations at small scales. Zamba 7B v1 was pre trained on 1T tokens of text and code data sourced from open web datasets. Subsequently in a second phase, Zamba was annealed on a mixture of 50B high quality tokens. Note: the current Huggingface implementation of Zamba performs slower than our internal implementation. We are working to fix this with the Huggingface team. Our technical report describing the training of Zamba is available here. Quick start Presequities To download Zamba, clone Zyphra's fork of transformers: 1. git clone https://github.com/Zyphra/transformers zamba 2. cd transformers zamba 3. Install the repository: pip install e . In order to run optimized Mamba implementations on a CUDA device, you need to install mamba ssm and causal conv1d : You can run the model without using the optimized Mamba kernels, but it is not recommended as it will result in significa…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy