Megatron LM & Megatron Core =========================== GPU optimized library for training transformer models at scale ⚡ Quick Start → Complete Installation Guide Docker, pip variants (dev,lts,etc.), and system requirements Latest News [2025/12] 🎉 Megatron Core development has moved to GitHub! All development and CI now happens in the open. We welcome community contributions. [2025/10] Megatron Dev Branch early access branch with experimental features. [2025/10] Megatron Bridge Bidirectional converter for interoperability between Hugging Face and Megatron checkpoints, featuring production ready recipes for popular models. [2025/08] MoE Q3 Q4 2025 Roadmap Comprehensive roadmap for MoE features including DeepSeek V3, Qwen3, advanced parallelism strategies, FP8 optimizations, and Blackwell performance enhancements. [2025/08] GPT OSS Model Advanced features including YaRN RoPE scaling, attention sinks, and custom activation functions are being integrated into Megatron Core. [2025/06] Megatron MoE Model Zoo Best practices and optimized configurations for training DeepSeek V3, Mixtral, and Qwen3 MoE models with performance benchmarking and checkpoint conversion tools. [2025/05] Megatron…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy