📕InternVideo2.5⚡ [\[📂 GitHub\]](https://github.com/OpenGVLab/InternVideo/tree/main/InternVideo2.5) [\[📜 Tech Report\]](https://arxiv.org/abs/2501.12386) InternVideo2.5 is a video multimodal large language model (MLLM, built upoon InternVL2.5) enhanced with long and rich context (LRC) modeling . It significantly improves upon existing MLLMs by enhancing their ability to perceive fine grained details and capture long form temporal structures. We achieve this through dense vision task annotations using direct preference optimization (TPO) and compact spatiotemporal representations via adaptive hierarchical token compression (HiCo). 📈 Performance VideoBenchmark Model MVBench LongVideoBench VideoMME(w/o sub) InternVideo2.5 75.7 60.6 65.1 Inference Speed We measured the average inference speed (tokens/s) of generating 1024 new tokens and 5198 (8192 2998) tokens with the context of an video (which takes 2998 tokens) under BF16 precision. w/ encoder indicates that the inference includes the time for video encoder. Quantization Speed (3022 tokens) Speed (8192 tokens) w/o encoder Speed(8192 tokens) w/ encoder BF16 33.40 31.91 21.33 INT4 31.95 26.37 The profiling runs on a single A800 SXM…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy