Dataset Summary Global MMLU 🌍 is a multilingual evaluation set spanning 42 languages, including English. This dataset combines machine translations for MMLU questions along with professional translations and crowd sourced post edits. It also includes cultural sensitivity annotations for a subset of the questions (2850 questions per language) and classifies them as Culturally Sensitive (CS) 🗽 or Culturally Agnostic (CA) ⚖️. These annotations were collected as part of an open science initiative led by Cohere Labs in collaboration with many external collaborators from both industry and academia. Curated by: Professional annotators and contributors of Cohere Labs Community. Language(s): 42 languages. License: Apache 2.0 Note: We also provide a "lite" version of Global MMLU called "Global MMLU Lite". This datatset is more balanced containing 200 samples each for CA and CS subsets for each language. And provides coverage for 15 languages with human translations. Global MMLU Dataset Family: Name Explanation Global MMLU Full Global MMLU set with translations for all 14K samples including CS and CA subsets Global MMLU Lite Lite version of Global MMLU with human translated samples in 15 la…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy