PIN 200M A mini version of "PIN: A Knowledge Intensive Dataset for Paired and Interleaved Multimodal Documents" Paper: https://arxiv.org/abs/2406.13923 This dataset contains around 200M samples in PIN format, with around 312 TB storage. 🚀 News [ 2025.09.22 ] !NEW! 🔥 We have completed the final version of the PIN 200M dataset and conducted some simple statistics on it. [ 2024.12.06 ] !NEW! 🔥 We have updated the quality signals, enabling a swift assessment of whether a sample meets the required specifications based on our quality indicators. Further detailed descriptions will be provided in the forthcoming formal publication. (Aside from the Chinese Markdown subset, there are unresolved issues that are currently being addressed.) This dataset contains 200M samples with PIN format. 0 Usage Download ALL files Download ONLY Jsonl files Decompression 1 Dataset statistics Subsect Documents ( ) Overall images ( ) Content images ( ) Documents (GB) Overall images (GB) Content images (GB) pg19 2,611,921 2,607,797 0 12.61 1,384.89 0.00 OBELICS 176,600,260 175,307,522 192,102,531 478.82 104,613.99 108,842.00 mmc4 core ff 9,425,497 9,332,329 15,732,454 59.98 5,579.99 9,659.09 chinese markdown…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy