PIN 14M A mini version of "PIN: A Knowledge Intensive Dataset for Paired and Interleaved Multimodal Documents" Paper: https://arxiv.org/abs/2406.13923 This dataset contains 14M samples in PIN format, with around 18.79 TB storage. 🚀 News [ 2025.09.04 ] !NEW! 🔥 We have completed the final version of the PIN 14M dataset and conducted some simple statistics on it. [ 2024.12.12 ] !NEW! 🔥 We have updated the quality signals for all subsets, with the dataset now containing 7.33B tokens after Llama3 tokenization. [ 2024.12.06 ] !NEW! 🔥 We have updated the quality signals, enabling a swift assessment of whether a sample meets the required specifications based on our quality indicators. Further detailed descriptions will be provided in the forthcoming formal publication. (Aside from the Chinese Markdown subset, there are unresolved issues that are currently being addressed.) This dataset contains 14M samples with PIN format. 0 Usage Download ALL files Download ONLY Jsonl files Decompression 1 Dataset statistics Subsect Documents ( ) Overall images ( ) Content images ( ) Documents (GB) Overall images (GB) Content images (GB) pg19 2,611,921 2,607,797 0 12.61 1,384.89 0.00 OBELICS 6,329,891…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy