Describe Anything: Detailed Localized Image and Video Captioning NVIDIA, UC Berkeley, UCSF Long Lian, Yifan Ding, Yunhao Ge, Sifei Liu, Hanzi Mao, Boyi Li, Marco Pavone, Ming Yu Liu, Trevor Darrell, Adam Yala, Yin Cui [Paper] [Code] [Project Page] [Video] [HuggingFace Demo] [Model/Benchmark/Datasets] [Citation] Model Card for DAM 3B Description Describe Anything Model 3B (DAM 3B) takes inputs of user specified regions in the form of points/boxes/scribbles/masks within images, and generates detailed localized descriptions of images. DAM integrates full image context with fine grained local details using a novel focal prompt and a localized vision backbone enhanced with gated cross attention. The model is for research and development only. This model is ready for non commercial use. License NVIDIA Noncommercial License Intended Usage This model is intended to demonstrate and facilitate the understanding and usage of the describe anything models. It should primarily be used for research and non commercial purposes. Model Architecture Architecture Type: Transformer Network Architecture: ViT and Llama This model was developed based on VILA 1.5. This model has 3B of model parameters. Inp…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy