Hy Embodied 0.5 VLA From Vision Language Action Models to a Real World Robot Learning Stack Tencent Robotics X × Tencent Hy Team 📖 Abstract We introduce Hy Embodied 0.5 VLA (Hy VLA) — an end to end Vision Language Action system that spans the full robot learning stack: data collection, model design, pre training, supervised fine tuning, RL post training, and real world deployment. Built on the Hy Embodied 0.5 MoT backbone, Hy VLA integrates a flow matching action expert, a compact memory encoder for multi frame history, and a delta chunk action representation decoupled from embodiment specific kinematics. Powered by 10,000+ hours of high fidelity UMI demonstrations collected via a custom fingertip interface with optical motion capture, Hy VLA achieves state of the art results on the RoboTwin 2.0 benchmark ( 90.9% / 90.1% on Clean / Randomized) and demonstrates robust cross embodiment transfer across four real world robot platforms. Paired with FlowPRO preference optimization and an asynchronous inference framework, Hy VLA establishes a scalable paradigm for continuous dexterous manipulation. Overview Hy Embodied 0.5 VLA Data is a large scale bimanual manipulation dataset for train…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy