Model Summary PowerMoE 3B is a 3B sparse Mixture of Experts (sMoE) language model trained with the Power learning rate scheduler. It sparsely activates 800M parameters for each token. It is trained on a mix of open source and proprietary datasets. PowerMoE 3B has shown promising results compared to other dense models with 2x activate parameters across various benchmarks, including natural language multi choices, code generation, and math reasoning. Paper: https://arxiv.org/abs/2408.13359 Usage Note: Requires installing HF transformers from source. Generation This is a simple example of how to use PowerMoE 3b model.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy