mmBERT Training Data (Ready to Use) Complete Training Dataset: Pre randomized and ready to use multilingual training data (3T tokens) for encoder model pre training. This dataset is part of the complete, pre shuffled training data used to train the mmBERT encoder models. Unlike the individual phase datasets, this version is ready for immediate use but the mixture cannot be modified easily. The data is provided in decompressed MDS format ready for use with ModernBERT's… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT pretraining data chunk0.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy