Model Details Model Card for Bamba 9B We introduce Bamba 9B, a decoder only language model based on the Mamba 2 architecture and is designed to handle a wide range of text generation tasks. It is trained from scratch using a two stage training approach. In the first stage, the model is trained on 2 trillion tokens from the Dolma v1.7 dataset. In the second stage, it undergoes additional training on 200 billion tokens, leveraging a carefully curated blend of high quality data to further refine its performance and enhance output quality. Model Params Layers Hidden Dim. Attention Heads GQA KV Heads Context Length Tied Embeddings Bamba 9B (9.78B) 32 4096 32 Yes 8 4096 False The current release includes the following models: Stage Bamba 9B Quantized Note Base Model ibm fms/Bamba 9B v1 ibm fms/Bamba 9B fp8 Stage 2 pretraining Base Model ibm fms/Bamba 9B 2T ibm fms/Bamba 9B fp8 Stage 1 pretraining Base Model ibm fms/Bamba 9B 1.8T ibm fms/Bamba 9B fp8 Intermediate checkpoints during Stage 1, more to come SFT coming soon coming soon to be released in the next drop DPO coming soon coming soon to be released in the next drop Original checkpoints (in dcp format) were also uploaded to public bu…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy