Introduction Reinforcement learning (RL) (e.g., GRPO) helps with grounding because of its inherent objective alignment—rewarding successful clicks—rather than encouraging long textual Chain of Thought (CoT) reasoning. Unlike approaches that rely heavily on verbose CoT reasoning, GRPO directly incentivizes actionable and grounded responses. Based on findings from our blog, we share state of the art GUI grounding models trained using GRPO. Grounding Performance We follow the standard evaluation protocol and benchmark our model on three challenging datasets. Our method consistently achieves the best results among all open source model families. Below are the comparative results: Model Size Open Source ScreenSpot V2 ScreenSpotPro OSWORLD G OSWORLD G Refined : : : : : : : : : : : : OpenAI CUA — ❌ 87.9 23.4 — — Claude 3.7 — ❌ 87.6 27.7 — — JEDI 7B 7B ✅ 91.7 39.5 54.1 — SE GUI 7B ✅ 90.3 47.0 — — UI TARS 7B ✅ 91.6 35.7 47.5 — UI TARS 1.5 7B ✅ 89.7 42.0 52.8 64.2 UGround v1 7B 7B ✅ — 31.1 — 36.4 Qwen2.5 VL 32B Instruct 32B ✅ 91.9 48.0 46.5 59.6 UGround v1 72B 72B ✅ — 34.5 — — Qwen2.5 VL 72B Instruct 72B ✅ 94.00 53.3 — 62.2 UI TARS 72B ✅ 90.3 38.1 — — OpenCUA 7B ✅ 92.3 50.0 55.3 68.3 OpenCUA…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy