🎬 Vript: Refine Video Captioning into Video Scripting [Github Repo] We construct a fine grained video text dataset with 12K annotated high resolution videos (~400k clips) . The annotation of this dataset is inspired by the video script. If we want to make a video, we have to first write a script to organize how to shoot the scenes in the videos. To shoot a scene, we need to decide the content, shot type (medium shot, close up, etc), and how the camera moves (panning, tilting, etc). Therefore, we extend video captioning to video scripting by annotating the videos in the format of video scripts. Different from the previous video text datasets, we densely annotate the entire videos without discarding any scenes and each scene has a caption with ~145 words. Besides the vision modality, we transcribe the voice over into text and put it along with the video title to give more background information for annotating the videos. Warning: Some zip files may contain empty folders. You can ignore them as these folders have no video clips and no annotation files. Getting Started By downloading these datasets, you agree to the terms of the License. The captions of the videos in the Vript dataset…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy