C Eval is a comprehensive Chinese evaluation suite for foundation models. It consists of 13948 multi choice questions spanning 52 diverse disciplines and four difficulty levels. Please visit our website and GitHub or check our paper for more details. Each subject consists of three splits: dev, val, and test. The dev set per subject consists of five exemplars with explanations for few shot evaluation. The val set is intended to be used for hyperparameter tuning. And the test set is for model evaluation. [2025.7.27] We have released the complete C Eval test set to the community! Now, you can directly evaluate on the C Eval test set more conveniently. Load the data More details on loading and using the data are at our github page. Please cite our paper if you use our dataset.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy