Model Summary Phi tiny MoE is a lightweight Mixture of Experts (MoE) model with 3.8B total parameters and 1.1B activated parameters. It is compressed and distilled from the base model shared by Phi 3.5 MoE and GRIN MoE using the SlimMoE approach, then post trained via supervised fine tuning and direct preference optimization for instruction following and safety. The model is trained on Phi 3 synthetic data and filtered public documents, with a focus on high quality, reasoning dense content. It is part of the SlimMoE series, which includes a larger variant, Phi mini MoE, with 7.6B total and 2.4B activated parameters. References: 📖 SlimMoE Paper 📖 Phi 3 Technical Report 📖 GRIN MoE Intended Uses Primary Use Cases The model is intended for commercial and research use in English. The model provides uses for general purpose AI systems and applications which require memory/compute constrained environments and latency bound scenarios. Use Case Considerations Our models are not specifically designed or evaluated for all downstream purposes. Developers should consider common limitations of language models as they select use cases, and evaluate and mitigate for accuracy, safety, and farine…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy