BLIP3o Pretrain Short Caption Dataset This collection contains 5 million images , each paired with a short (~20 token) caption generated by Qwen/Qwen2.5 VL 7B Instruct . Download Load Dataset without Extracting You don’t need to unpack the .tar archives, use WebDataset support in 🤗datasets instead: Feel free to comment on it when you have any issues.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy