MzansiLM 125M MzansiLM is a 125M parameter decoder only language model trained from scratch on MzansiText , a multilingual corpus covering all eleven official South African languages. Model Details Parameters: 125,008,384 Architecture: decoder only LlamaForCausalLM Hidden size: 512 Intermediate size: 1536 Layers: 30 Attention heads: 9 Key/value heads: 3 Context length: 2048 RoPE theta: 10000.0 RMSNorm epsilon: 1e 5 Tied word embeddings: true Training attention implementation: flash attention 2 Tokenizer MzansiLM uses a custom BPE tokenizer with a vocabulary size of 65536 . [BOS] = 0 [EOS] = 1 [PAD] = 2 [UNK] = 3 Normalizer: NFD Pre tokenizer: ByteLevel Post processing: single sequence: [BOS] $A [EOS] pair sequence: [BOS] $A [EOS] [BOS] $B [EOS] Training Data The model was trained on MzansiText and covers all eleven official South African languages: af , en , nso , sot , ssw , tsn , tso , ven , xho , zul , nbl Related releases: Paper: arXiv:2603.20732 Raw corpus: anrilombard/mzansi text Tokenized corpus: anrilombard/mzansi text tokenized GitHub code and configs: https://github.com/Anri Lombard/sallm Intended Use MzansiLM is a research model for pretraining, fine tuning, and evaluati…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy