VisionArena Battle: 30K Real World Image Conversations with Pairwise Preference Votes 200k single and multi turn chats between users and VLM's collected on Chatbot Arena. WARNING: Images may contain inappropriate content. Dataset Details 200K conversations 45 VLM's 138 languages ~43k unique images Question Category Tags (Captioning, OCR, Entity Recognition, Coding, Homework, Diagram, Humor, Creative Writing, Refusal) Dataset Description 200,000 conversations where users interact with two anonymized VLMs,collected through the open source platform Chatbot Arena, where users chat with LLMs and VLMs through direct chat, side by side, or anonymous side by side chats. Users provide preference votes for responses, which are aggregated using the Bradley Terry model to compute leaderboard rankings. Data for anonymous side by side chats can be found here. The dataset includes conversations from February 2024 to September 2024. Users explicitly agree to have their conversations shared before chatting. We apply an NSFW, CSAM, PII (text, and face detectors (1, 2) to remove any inappropriate images, personally identifiable images/text, or images with human faces. These detectors are not perfect,…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy