MMLU ProX Multilingual Model Predictions Raw per sample model predictions on MMLU ProX across 29 languages and 25 open weight LLMs , produced with lm evaluation harness. This dataset releases the full prediction logs (not just aggregate scores) so that item level responses can be re analysed — e.g. for Item Response Theory (IRT) modelling of multilingual benchmarks, error analysis, or per item difficulty estimation. Repository structure Languages (29): af, ar, bn, cs, de, en, es, fr, hi, hu, id, it, ja, ko, mr, ne, pt, ru, sr, sw, te, th, uk, ur, vi, wo, yo, zh, zu Subjects (14): biology, business, chemistry, computer science, economics, engineering, health, history, law, math, other, philosophy, physics, psychology Models (25): CohereLabs aya expanse (8b/32b) & tiny aya global; Qwen3.5 (2B/4B/9B/27B + 35B A3B/122B A10B MoE); allenai Olmo 3 7B & Olmo 3.1 32B (Instruct/Think); google gemma 3 (4b/12b/27b) it; meta llama Llama 3.1 8B / 3.2 3B / 3.3 70B Instruct; microsoft Phi 4 mini (instruct/reasoning) & Phi 4 reasoning( plus); openai gpt oss (20b/120b) File schemas samples .jsonl — one JSON object per evaluated item Field Description doc id Index of the item within the subject split…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy