RDT2 VQ: Vision Language Action with Residual VQ Action Tokens RDT2 VQ is an autoregressive Vision Language Action (VLA) model adapted from Qwen2.5 VL 7B Instruct and trained on large scale UMI bimanual manipulation data. It predicts a short horizon relative action chunk (24 steps, 20 dims/step) from binocular wrist camera RGB and a natural language instruction. Actions are discretized with a lightweight Residual VQ (RVQ) tokenizer, enabling robust zero shot transfer across unseen embodiments for simple, open vocabulary skills (e.g., pick, place, shake, wipe). Home Github Discord Paper Table of contents Highlights Model details Hardware & software requirements Quickstart (inference) Precision settings Intended uses & limitations Troubleshooting Changelog Citation Contact Highlights Zero shot cross embodiment : Demonstrated on Bimanual UR5e and Franka Research 3 setups; designed to generalize further with correct hardware calibration. UMI scale : Trained on 10k+ hours from 100+ indoor scenes of human manipulation with the UMI gripper. Residual VQ action tokenizer : Compact, stable action codes; open vocabulary instruction following via Qwen2.5 VL 7B backbone. Model details Architect…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy