Overview This dataset contains user votes collected in the text only category. Each row represents a single vote judging two models (model a and model b) on a user conversation, along with the full conversation history and metadata. Key fields include: id : Unique feedback ID of each vote/row. evaluation session id : Unique ID of each evaluation session, which can contain multiple separate votes/evaluations. evaluation order : Evaluation order of the current vote. winner : Battle result containing either model a, model b, tie, or both bad. conversation a/conversation b : Full conversation of the current evaluation order. full conversation : The entire conversation, including context prompts and answers from all previous evaluation orders. Note that after each vote new models are sampled, thus the responding models vary across the full context. conv metadata : Aggregated markdown and token counts for style control. category tag : Annotation tags including the categories math, creative writing, hard prompts, and instruction following. is code : Whether the conversation involves code. License User prompts are licensed under CC BY 4.0, and model outputs are governed by the terms of use…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy