We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy
mmBERT Training Data (Ready-to-Use) Complete Training Dataset: Pre-randomized and ready-to-use multilingual training data (3T tokens) for encoder model pre-training. This dataset is part of the complete, pre-shuffled training data used to train the mmBERT encoder models. Unlike the individual phase datasets, this version is ready for immediate use but the mixture cannot be modified easily. The data is provided in decompressed MDS format ready for use with ModernBERT's… See the full description on the dataset page: https://huggingface.co/datasets/orionweller/mmBERT-data-midtraining.
No dataset card provided yet.
Mirrored from an external registry.
Last synced 6/20/2026
Preview not yet available for this dataset