Polysemous Words Polysemous Words is a large scale collection of contextual examples for 200 common polysemous English words. Each word in this dataset has multiple, distinct senses and appears across a wide variety of natural web text contexts, making the dataset ideal for research in word sense induction (WSI) , word sense disambiguation (WSD) , and probing the contextual understanding capabilities of large language models (LLMs) . Dataset Overview This dataset is built from the 350B variant of the FineWeb Edu corpus, a large scale, high quality web crawl dataset recently used in training state of the art LLMs. We focus on 200 polysemous English words selected for their frequency and semantic diversity. The initial list was generated using ChatGPT, then refined through manual inspection to ensure each word has multiple distinct meanings. For each target word: The FineWeb Edu corpus was filtered to retain only samples containing a surface form match (the full word without regards to capitalization). The final dataset contains over 470 million contextual examples, totaling about 3 TB of raw text. Data is organized by word, allowing easy per word processing or analysis. 📍 Dataset U…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy