RLDX 1 VLM Paper · Project page · Code · Models RLDX 1 VLM is the vision language backbone used by the RLDX 1 robot policy family. It is a Qwen3 VL 8B Instruct checkpoint distributed separately from the action policy so that researchers can inspect, finetune, or replace the perceptual stack independently. Note. This checkpoint exposes a standard Qwen3 VL VLM interface only — it does not ship the Multi Stream Action Transformer head, the cognition tokens, the memory / motion / physics modules, or the RLDX inference server. For action prediction, use one of the RLDX 1 PT , RLDX 1 FT , or RLDX 1 MT checkpoints. Intended use As the backbone path for finetuning a fresh RLDX 1 policy (recipe). For VLM only ablations, dense captioning experiments, or perceptual probing within the RLDX research stack. Quick start Model details Type: Vision language model (multimodal text + image / video). Backbone: Qwen/Qwen3 VL 8B Instruct . Params: 8B. Role in RLDX 1: perceptual encoder for the MSAT action policy. Cognition tokens are injected into this backbone and routed through Qwen3 VL hidden states to produce a compact perceptual summary consumed by the action mod…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy