TVP base model The TVP model was proposed in Text Visual Prompting for Efficient 2D Temporal Video Grounding by Yimeng Zhang, Xin Chen, Jinghan Jia, Sijia Liu, Ke Ding. The goal of this model is to incorporate trainable prompts into both visual inputs and textual features to temporal video grounding(TVG) problems. It was introduced in this paper. Model Detail Description Model Authors Yimeng Zhang, Xin Chen, Jinghan Jia, Sijia Liu, Ke Ding Date 2023 Version Base Type Text Visual Prompting for Temporal Video Grounding Paper or Other Resources Paper: Text Visual Prompting for Efficient 2D Temporal Video Grounding Dataset: Charades License Other Questions or Comments Community Tab and Intel DevHub Discord Intended Use Description Primary intended uses The TVP model is designed for temporal video grounding (TVG), specifically to predict the start and end times of moments described by a text sentence within a long, untrimmed video. Primary intended users Researchers and developers working in the field of computer vision, particularly those focused on video understanding and cross modal (text and video) tasks. Out of scope uses The model is not intended for real time video processing or…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy