Dataset Card for WildChat Dataset Description Paper: https://arxiv.org/abs/2405.01470 Interactive Search Tool: https://wildvisualizer.com (paper) License: ODC BY Language(s) (NLP): multi lingual Point of Contact: Yuntian Deng Dataset Summary WildChat is a collection of 1 million conversations between human users and ChatGPT, alongside demographic data, including state, country, hashed IP addresses, and request headers. We collected WildChat by offering online users free access to OpenAI's GPT 3.5 and GPT 4. In this version, 25.53% of the conversations come from the GPT 4 chatbot, while the rest come from the GPT 3.5 chatbot. The dataset contains a broad spectrum of user chatbot interactions that are not previously covered by other instruction fine tuning datasets: for example, interactions include ambiguous user requests, code switching, topic switching, political discussions, etc. WildChat can serve both as a dataset for instructional fine tuning and as a valuable resource for studying user behaviors. Note that this version of the dataset only contains non toxic user inputs/ChatGPT responses. Updates 2024 10 17: Content Update. Conversations flagged by Niloofar Mireshghallah and h…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy