Model Overview Description: Cosmos Embed1 is a joint video text embedder tailored for physical AI. It can be used for text to video retrieval, inverse video search, semantic deduplication, zero shot and k nearest neighbors (kNN) classification, and as a base model for video curation tasks. It has state of the art (SOTA) performance on autonomous vehicle (AV) and robotics datasets, while maintaining competitive performance in general domains. A fine tuned variant is also provided for video anomaly detection and classification. This model is ready for commercial use. The Cosmos Embed1 release includes the following embedders: Variant Resolution Frames Embedding Dim Cosmos Embed1 224p 224×224 8 256 Cosmos Embed1 336p 336×336 8 768 Cosmos Embed1 448p 448×448 8 768 Note: while each checkpoint was optimized at a specific fixed resolution (and default to these), they all support arbitrary non square resolutions. In addition, a fine tuned variant is provided for anomaly detection applications: Variant Base Model Resolution Frames Fine tuning Dataset Embedding Dim Cosmos Embed1 448p anomaly detection Cosmos Embed1 448p 448×448 8 Vad Reasoning (training set) 768 The Cosmos Embed1 448p anomal…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy