Cosmos Embed1 : A joint video text embedder for physical AI Website Hugging Face Demo app Model Overview Description Cosmos Embed1 is a joint video text embedder tailored for physical AI. It can be used for text to video retrieval, inverse video search, semantic deduplication, zero shot and k nearest neighbors (kNN) classification, and as a base model for video curation tasks. It has state of the art (SOTA) performance on autonomous vehicle (AV) and robotics datasets, while maintaining competitive performance in general domains. This model is ready for commercial use. Model Developer : NVIDIA Model Versions The Cosmos Embed1 release includes the following embedders: Cosmos Embed1 Cosmos Embed1 224p (optimized with 8 frames and 224x224 input resolution, 256 dim output text and video embeddings) Cosmos Embed1 336p (optimized with 8 frames and 336x336 input resolution, 768 dim output text and video embeddings) Cosmos Embed1 448p (optimized with 8 frames and 448x448 input resolution, 768 dim output text and video embeddings) Note, while each checkpoint was optimized at a specific fixed resolution (and default to these), they all support arbitrary non square resolutions. License This mo…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy