LocateAnything Data 中文 · Paper · Model · Code Overview LocateAnything Data is the public training data release for LocateAnything: Fast and High Quality Vision Language Grounding with Parallel Box Decoding . LocateAnything formulates detection and visual grounding as a unified vision language task. Given an image and a category, phrase, text string, or action oriented instruction, the model predicts the corresponding bounding boxes or points. The data spans natural images, dense scenes, people, autonomous driving, embodied interaction, graphical user interfaces, scene text, documents, and tables. The paper's Parallel Box Decoding treats a box or point as a structured geometric unit instead of generating its coordinates independently. LocateAnything Data provides the diverse spatial supervision used to train this unified formulation across visual domains. This repository provides: detection, grounding, and pointing annotations in previewable JSONL; image media packed as indexed WebDataset TAR shards; Megatron Energon metadata for distributed training; and public mappings from every training record and packed image back to its original dataset relative media name. Data coverage The f…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy