Dataset Description Repository: openai/gpt2 Paper: Radford et al. Language Models are Unsupervised Multitask Learners Dataset Summary This dataset is comprised of the LAMBADA test split as pre processed by OpenAI (see relevant discussions here and here). It also contains machine translated versions of the split in German, Spanish, French, and Italian. LAMBADA is used to evaluate the capabilities of computational models for text understanding by means of a word prediction task. LAMBADA is a collection of narrative texts sharing the characteristic that human subjects are able to guess their last word if they are exposed to the whole text, but not if they only see the last sentence preceding the target word. To succeed on LAMBADA, computational models cannot simply rely on local context, but must be able to keep track of information in the broader discourse. Languages English, German, Spanish, French, and Italian. Source Data For non English languages, the data splits were produced by Google Translate. See the translation script.py for more details. Additional Information Hash Checksums For data integrity checks we leave the following checksums for the files in this dataset: File Name…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy