FineWeb Mask 📜 DATAMASK Paper 💻 GitHub Repository 📦 Fineweb Mask Dataset 📚 Introduction FineWeb Mask is a 1.5 trillion token, high efficiency pre training dataset curated using the DATAMASK framework. Developed by the ByteDance Seed team , DATAMASK addresses the fundamental tension in large scale data selection: the trade off between high quality and high diversity . By modeling data selection as a Mask Learning problem, we provide a derivative of the original FineWeb corpus. FineWeb Mask is designed to eliminate semantic redundancy while preserving the highest quality samples, allowing models to achieve superior performance with significantly less data. 🎯 The Problem: The Quality Diversity Trap In large language model (LLM) pre training, developers usually face two suboptimal choices: 1. The Quality Trap: Filtering solely by quality scores leads to "diminishing returns." Samples become highly clustered, resulting in severe semantic redundancy. 2. The Diversity Trap: Filtering solely for diversity often discards high value quality samples, leading to worse performance than the original raw dataset. 3. The Compute Bottleneck: Traditional diversity algorithms (like greedy select…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy