The visual encoder of VideoLLaMA 3: Frontier Multimodal Foundation Models for Video Understanding If you like our project, please give us a star ⭐ on Github for the latest update. 🌟 Introduction This model serves as the visual encoder in VideoLLaMA3. VideoLLaMA3 leverages the Any resolution Vision Tokenization (AVT) approach to dynamically process images and videos of varying resolutions. This is accomplished by adapting the pre trained vision encoder (based on ViT architecture) to use 2D RoPE (Rotary Position Embeddings), replacing the absolute position embeddings traditionally used in ViT. With AVT, VideoLLaMA3 is able to represent images and videos with greater detail across different resolutions, enriching the vision tokens with more information. To ensure seamless integration with AVT, we fine tune both the vision encoder and the projector during the Vision Encoder Adaptation stage (Stage 1 in the VideoLLaMA3 training pipeline) using scene images, document data, and scene images with text. Before training, the model parameters and architecture are initialized from SigLip. 🚀 Model Porfermance Base Model GQA AI2D ChartQA DocVQA val MME clip vit large patch14 336 61.50 56.28 18…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy