CIA Declassified Reading Room HF Library Target account: manus4oHER This project is a streaming pipeline for building a Hugging Face dataset mirror of public CIA declassified Reading Room / CREST records without staging the full corpus on this laptop. The laptop stores only scripts, small manifests, and logs. Bulk crawling should run in Hugging Face Jobs, one bounded page range per job. Each job uploads its own shard and then exits. Dataset Shape metadata/ / /documents.jsonl.gz texts/ / /texts.jsonl.gz pdf shards/ / /pdfs.zip when PDF capture is enabled reports/ / /crawl report.json manifests/source config.json scripts/ for reproducible crawling Source Scope Initial source target: CIA Electronic Reading Room / CREST public records. Internet Archive mirrors of CIA CREST / CIA Reading Room records when the live CIA route loops or blocks programmatic access. The CIA endpoint currently redirect loops from this laptop for naive HTTP clients and also looped from the first HF smoke job. The pipeline therefore supports Internet Archive as a practical source of the same public domain CIA document corpus while preserving original CIA document identifiers and source provenance. Operating Rule…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy