VideoPrism Model Card Paper : https://huggingface.co/papers/2402.13217 arXiv : https://arxiv.org/pdf/2402.13217 GitHub : https://github.com/google deepmind/videoprism Blog : https://research.google/blog/videoprism a foundational visual encoder for video understanding/ VideoPrism is a foundational video encoder that enables state of the art performance on a large variety of video understanding tasks. It takes video frames as input and outputs compact embeddings of the frames, which one can conveniently feed into classifiers, LLMs, retrieval models, etc. When tested on 33 public video understanding benchmarks over four task categories, a single frozen VideoPrism checkpoint outperforms previous best performing foundation models on 31 of them, with no fine tuning on target task datasets. Model details We release the following model variants: Model Name Configuration Name Model Type Backbone Params File Size Checkpoint : : : : : : : : VideoPrism B videoprism public v1 base Video encoder ViT B 114M 458MB link VideoPrism L videoprism public v1 large Video encoder ViT L 354M 1.42GB link VideoPrism LvT B videoprism lvt public v1 base Video text encoders ViT B 248M 991MB link VideoPrism LvT…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy