RoboOmni: Proactive Robot Manipulation in Omni modal Context 📖 arXiv Paper (Accepted to ICLR 2026 🎉) 🌐 Website 🤗 Model 🤗 Dataset 🛠️ Github Recent advances in Multimodal Large Language Models (MLLMs) have driven rapid progress in Vision–Language–Action (VLA) models for robotic manipulation. Although effective in many scenarios, current approaches largely rely on explicit instructions, whereas in real world interactions, humans rarely issue… See the full description on the dataset page: https://huggingface.co/datasets/Whoisjutanlee/2.1tbofdata.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy