LocateAnything: Fast and High Quality Vision Language Grounding with Parallel Box Decoding ๐ Quick Links ๐ Online Demo : LocateAnything (Hugging Face Spaces) ๐ป GitHub Code : NVlabs/Eagle/Embodied ๐ Paper : arXiv:2605.27365 Model Overview Description: LocateAnything is a vision language model for fast and high quality visual grounding, enabling precise object localization, dense detection, and point based localization across diverse domains in both Enterprise Intelligence and Physical AI. The model adopts a generalist design, supporting tasks such as referring expression grounding, multi object detection, GUI element grounding, and text localization, with strong performance in complex and cluttered scenes. Its core innovation, Parallel Box Decoding (PBD), predicts complete bounding box coordinates in a single parallel step rather than autoregressive token by token decoding, improving efficiency while preserving geometric consistency. This enables up to 2.5ร higher throughput compared to prior approaches. The model is trained on a large scale multi domain dataset (12M images, 138M+ queries, 785M bounding boxes) spanning natural scenes, robotics, driving, GUI interaction, and docuโฆ
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy