dolma3-6t-sample-100000-docs
Materialized stratified sample of 100,000 docs per bin (~58k source-shard files, ~186 GB) from the deduplicated 6T Dolma3 corpus. Seed 42, materialized by 128 Modal workers (worker_0000/ through worker_0127/).
Layout
HCAI-Lab/dolma3-6t-sample-100000-docs/
├── bin_summary.csv
├── sample_contract.json
└── worker_NNNN/
└── soc127__phase1_pool_shared__...__shard_NNNNNNNN.jsonl.zst (~450 files per worker)
Dual-access companion
This dataset is the dataset-API twin of the existing bucket:
- Bucket:
hf://buckets/HCAI-Lab/dolma3-6t-sample-100000-docs(S3-style access) - Dataset:
HCAI-Lab/dolma3-6t-sample-100000-docs(this repo;snapshot_download+ dataset viewer)
Both surfaces contain bit-identical data. Pick whichever access pattern fits your tooling.
Provenance
Created 2026-05-25 as part of the HCAI-Lab HF org cleanup (PR 3): every sample size now has both a bucket and a dataset twin for symmetric access. The other sample sizes (dolma3-6t-sample-{500,1000,5000,10000,50000}-docs) follow the same pattern.
| Field | Value |
|---|---|
| Original bucket | HCAI-Lab/dolma3-6t-sample-100000-docs (created 2026-03-24) |
| Dataset twin created | 2026-05-25 |
| Source files | 58,264 (bit-exact match to bucket) |
| Size | ~184 GB on HF dataset accounting, ~186.7 GB on bucket accounting (~1.3% diff is HF-side compression variance) |
See docs/data_home/inventory.json and docs/HCAI_LAB_NAMING_CONVENTION.md for the full org-wide inventory and naming rule.