ShareChat: A Dataset of Chatbot Conversations in the Wild This dataset contains 142,808 real world user conversations across multiple conversational AI platforms (ChatGPT, Claude, Gemini, Grok, and Perplexity). The dataset is collected and processed for research purposes to understand usage patterns, topic distributions, and behavioral characteristics across different AI platforms. Basic Statistics Metric Value Total Conversations 142,808 Total Turns 660,293 Average Turns per Conversation 4.62 Languages Covered 101 Collection Period April 2023 – October 2025 Avg. User Tokens 135.04 ± 1,820.88 Avg. Chatbot Tokens 1,115.30 ± 1,764.81 Per Platform Distribution Platform Conversations Turns Avg. Turns Languages ChatGPT 102,740 542,148 5.28 101 Perplexity 17,305 24,378 1.41 45 Grok 14,415 53,094 3.69 60 Gemini 7,402 36,422 4.92 47 Claude 946 4,251 4.49 19 Data Structure These released DataFrames provide turn level conversation records from five platforms with a shared core schema, where each row is one message. All datasets include: platform , url , turns count , message index, role, plain text , and detected language final , enabling consistent cross platform analysis of conversation st…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy