ULVR-filtered
Filtered subset of RuoliuYang/ULVR_v2_clean: the 101,951 training samples that Qwen2.5-VL-7B-Instruct answered incorrectly given only input_image, but correctly once the intermediate_image_* were also provided (judged by Qwen3-VL-32B-Instruct). Same schema / subsets / train-split structure as the source.
| subset | rows |
|---|---|
| scene_graph | 3522 |
| edge | 1394 |
| depth | 537 |
| segmentation | 1328 |
| bbox_highlight | 15186 |
| bbox_crop | 15260 |
| text_cot | 27158 |
| helper_interleaved | 37566 |
| total | 101951 |