cua lite/Aguvis cua lite preprocessed version of Aguvis (xlangai/aguvis stage1 + xlangai/aguvis stage2). A large scale composite GUI dataset using a unified PyAutoGUI action format across mobile / web / desktop. Stage 1 → grounding.action (single step locate and click across SeeClick, GUIEnv, WebUI, RicoSCA, RICO Icon, Widget Captioning, UI RefExp, OmniACT); Stage 2 → navigation (multi step trajectories from AndroidControl, AitW, MiniWoB++, COAT, GUIDE). Origin https://huggingface.co/datasets/xlangai/aguvis stage1 https://huggingface.co/datasets/xlangai/aguvis stage2 Load via datasets You can also filter by metadata.platform / metadata.task type / metadata.others. after loading; every row carries a rich metadata struct (see schema below). Schema Each row has these columns: column type notes images list[Image] embedded PNG/JPEG bytes; HF viewer renders thumbnails messages list[struct] OpenAI style turns with role + structured content metadata struct {platform, task type, extra tool schemas, valid actions, others{...}} Coordinate values in messages are normalized to [0, 1000] integers. Image dedup ( grounding. / understanding cohorts). These cohorts are single image per row and many…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy