WeChat MiniMax VL 01 1. Introduction We are delighted to introduce our MiniMax VL 01 model. It adopts the "ViT MLP LLM" framework, which is a commonly used technique in the field of multimodal large language models. The model is initialized and trained with three key parts: a 303 million parameter Vision Transformer (ViT) for visual encoding, a randomly initialized two layer MLP projector for image adaptation, and the MiniMax Text 01 as the base LLM. MiniMax VL 01 has a notable dynamic resolution feature. Input images are resized per a pre set grid, with resolutions from 336×336 to 2016×2016, keeping a 336×336 thumbnail. The resized images are split into non overlapping patches of the same size. These patches and the thumbnail are encoded separately and then combined for a full image representation. The training data for MiniMax VL 01 consists of caption, description, and instruction data. The Vision Transformer (ViT) is trained on 694 million image caption pairs from scratch. Across four distinct stages of the training pipeline, a total of 512 billion tokens are processed, leveraging this vast amount of data to endow the model with strong capabilities. Finally, MiniMax VL 01 has r…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy