Two Box Judge GUI Dataset A multimodal dataset for training GUI element selection models. Given two candidate bounding boxes on a GUI screenshot, the model learns to select the one that better fulfills the user's intent. Dataset Description This dataset is designed for training judge models in GUI grounding pipelines. When a visual grounding model produces multiple candidate regions, the judge model determines which candidate best matches the user's command. Key Statistics Metric Value Total Samples 128,487 (balanced) Training Samples 115,638 Validation Samples 12,849 Image Pairs ~257K (2 images per sample) Label Distribution 92% bbox1, 8% bbox2 Data Sources GUIAct Dataset : Web (single/multi step) interactions (~70K samples) Desktop Domain Dataset : Windows/Mac desktop applications (~423K samples) Usage Loading with Datasets Library Loading Locally Data Format Each sample contains: Field Type Description sample id int Unique sample identifier image1 string Path to image with green box (candidate 1) image2 string Path to image with red box (candidate 2) user command string Natural language instruction label string Correct answer: "1" or "2" Sample Example Image Description Image 1…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy