MC Math Rollouts Step level Monte Carlo resampling rollouts of Qwen3 models on competition math benchmarks, for measuring the advantage and importance of individual reasoning steps . 🔎 Explore the data interactively: MC Math Rollouts Value Profiles viewer — browse per step value profiles for individual traces without downloading anything. For each seed response, the chain of thought is split into steps ("reasoning prefixes"). From the end of every prefix, the model is resampled ~50 times to completion, and the resulting final answer distributions are compared across consecutive prefixes. This yields per step estimates of how much each reasoning step changes the probability of reaching the correct answer, alongside token level entropy and LLM judged semantic labels for every step. Note: This is a raw experiment artifact repository (~3.15 TB of nested CSV/JSON files), not a load dataset ready dataset, so the built in dataset viewer is disabled — use the viewer Space instead. See Usage for how to download individual files. Benchmarks and models Benchmark Prompts Config prefix GSM8K 30 gsm8k mc10 MATH 500 30 math500 mc10 AIME 2024 30 aime24 mc10 AIME 2025 30 aime25 mc10 AIME 2026 30 a…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy