ERNIE 4.5 21B [!NOTE] Note: " Paddle " models use PaddlePaddle weights, while " PT " models use Transformer style PyTorch weights. ERNIE 4.5 Highlights The advanced capabilities of the ERNIE 4.5 models, particularly the MoE based A47B and A3B series, are underpinned by several key technical innovations: 1. Multimodal Heterogeneous MoE Pre Training: Our models are jointly trained on both textual and visual modalities to better capture the nuances of multimodal information and improve performance on tasks involving text understanding and generation, image understanding, and cross modal reasoning. To achieve this without one modality hindering the learning of another, we designed a heterogeneous MoE structure , incorporated modality isolated routing , and employed router orthogonal loss and multimodal token balanced loss . These architectural choices ensure that both modalities are effectively represented, allowing for mutual reinforcement during training. 2. Scaling Efficient Infrastructure: We propose a novel heterogeneous hybrid parallelism and hierarchical load balancing strategy for efficient training of ERNIE 4.5 models. By using intra node expert parallelism, memory efficient p…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy