VideoLLaMA 3: Frontier Multimodal Foundation Models for Video Understanding If you like our project, please give us a star ⭐ on Github for the latest update. 📰 News [2024.01.24] 🔥🔥 Online Demo is available: VideoLLaMA3 Image 7B, VideoLLaMA3 7B. [2024.01.22] Release models and inference code of VideoLLaMA 3. 🌟 Introduction VideoLLaMA 3 represents a state of the art series of multimodal foundation models designed to excel in both image and video understanding tasks. Leveraging advanced architectures, VideoLLaMA 3 demonstrates exceptional capabilities in processing and interpreting visual content across various contexts. These models are specifically designed to address complex multimodal challenges, such as integrating textual and visual information, extracting insights from sequential video data, and performing high level reasoning over both dynamic and static visual scenes. 🌎 Model Zoo Model Base Model HF Link VideoLLaMA3 7B ( This Checkpoint ) Qwen2.5 7B DAMO NLP SG/VideoLLaMA3 7B VideoLLaMA3 2B Qwen2.5 1.5B DAMO NLP SG/VideoLLaMA3 2B VideoLLaMA3 7B Image Qwen2.5 7B DAMO NLP SG/VideoLLaMA3 7B Image VideoLLaMA3 2B Image Qwen2.5 1.5B DAMO NLP SG/VideoLLaMA3 2B Image We also upl…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy