MMLU ProX Multilingual Model Predictions Raw per sample model predictions on MMLU ProX across 29 languages and 25 open weight LLMs , produced with lm evaluation harness. This dataset releases the full prediction logs (not just aggregate scores) so that item level responses can be re analysed — e.g. for Item Response Theory (IRT) modelling of multilingual benchmarks, error analysis, or per item difficulty estimation. Repository structure Languages (29): af, ar, bn, cs, de, en, es, fr, hi, hu, id, it, ja, ko, mr, ne, pt, ru, sr, sw, te, th, uk, ur, vi, wo, yo, zh, zu Subjects (14): biology, business, chemistry, computer science, economics, engineering, health, history, law, math, other, philosophy, physics, psychology Models (25): CohereLabs aya expanse (8b/32b) & tiny aya global; Qwen3.5 (2B/4B/9B/27B + 35B A3B/122B A10B MoE); allenai Olmo 3 7B & Olmo 3.1 32B (Instruct/Think); google gemma 3 (4b/12b/27b) it; meta llama Llama 3.1 8B / 3.2 3B / 3.3 70B Instruct; microsoft Phi 4 mini (instruct/reasoning) & Phi 4 reasoning( plus); openai gpt oss (20b/120b) File schemas samples .jsonl — one JSON object per evaluated item Field Description doc id Index of the item within the subject split…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy