BigScience BLOOM Evaluation Results This repository contains evaluation results & original predictions of BLOOM & friends. Usage You can load numeric results via: If it takes too long, it may be faster to clone the repository and load the data from disk: For example generations (.jsonl files), you need to manually browse the repository. Structure For bigsciencelmevalharness , lmevalharness & codeeval evaluation frameworks the structure is: model name evaluation framework checkpoint type dataset name data Evaluation Procedure bigsciencelmevalharness files were created using the below: https://github.com/bigscience workshop/Megatron DeepSpeed/pull/291 https://github.com/bigscience workshop/lm evaluation harness lmevalharness files were created using the below: https://github.com/bigscience workshop/Megatron DeepSpeed https://github.com/EleutherAI/lm evaluation harness codeeval files were created using the HumanEval code dataset with the below: https://github.com/loubnabnl/bloom code evaluation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy