Legal Corpus Raw Batches This repository stores raw, source preserving legal domain corpus batches collected for legal language model pretraining, retrieval, embedding, and corpus analysis work. It is intentionally batch oriented: each folder corresponds to one source slice, shard, or non overlapping range, with source metadata and upload verification artifacts kept alongside the raw files. Last draft card update: 2026 06 30 09:54 UTC . Current Build Status Accounted uncompressed text bytes: 2,202,530,242,387 bytes (~ 2.203 TB ). Batch entries tracked in progress state: 134 . Verified uploaded batch IDs in local ledger: 134 . Source descriptors in provenance audit: 225 . Target collection size: at least 7 TB, preferably 10 TB, of useful raw legal text with provenance. Collection is ongoing; numbers above are generated from cluster state files and may lag active downloads until scan/upload verification completes. What Is Included The corpus is focused on legal and law adjacent primary or research useful text: legislation, regulations, public laws, court decisions, administrative materials, legal filings, contracts, patents/regulatory IP material, and multilingual/parallel legal corp…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy