Fine Vision FineVision is a massive collection of datasets with 17.3M images , 24.3M samples , 88.9M turns , and 9.5B answer tokens , designed for training state of the art open Vision Language Models. More detail can be found in the blog post: https://huggingface.co/spaces/HuggingFaceM4/FineVision The version in this repository concatenated all the configs in the original dataset and then shuffled them. This is done to facilitate streaming the data directly from the hub! Load the data Structure Categories Licensing Information Each of the publicly available sub datasets present in FineVision are governed by specific licensing conditions. Therefore, when making use of them you must take into consideration each of the licenses governing each dataset. To the extent we have any rights in the prompts, these are licensed under CC BY 4.0. Citation If you find this dataset useful, please cite:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy