Clone from "friedrichor/MSR VTT". MSRVTT contains 10K video clips and 200K captions. We adopt the standard 1K A split protocol, which was introduced in JSFusion and has since become the de facto benchmark split in the Text Video Retrieval field. Train: train 7k: 7,010 videos, 140,200 captions train 9k: 9,000 videos, 180,000 captions Test: test 1k: 1,000 videos, 1,000 captions 🌟 Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy