dolma3 6t sample 100000 docs Materialized stratified sample of 100,000 docs per bin (~58k source shard files, ~186 GB) from the deduplicated 6T Dolma3 corpus. Seed 42, materialized by 128 Modal workers ( worker 0000/ through worker 0127/ ). Layout Dual access companion This dataset is the dataset API twin of the existing bucket: Bucket : hf://buckets/HCAI Lab/dolma3 6t sample 100000 docs (S3 style access) Dataset : HCAI Lab/dolma3 6t sample 100000 docs (this repo; snapshot download + dataset viewer) Both surfaces contain bit identical data. Pick whichever access pattern fits your tooling. Provenance Created 2026 05 25 as part of the HCAI Lab HF org cleanup (PR 3): every sample size now has both a bucket and a dataset twin for symmetric access. The other sample sizes ( dolma3 6t sample {500,1000,5000,10000,50000} docs ) follow the same pattern. Field Value Original bucket HCAI Lab/dolma3 6t sample 100000 docs (created 2026 03 24) Dataset twin created 2026 05 25 Source files 58,264 (bit exact match to bucket) Size ~184 GB on HF dataset accounting, ~186.7 GB on bucket accounting (~1.3% diff is HF side compression variance) See docs/data home/inventory.json and docs/HCAI LAB NAMING CON…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy