We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy
CatholicCorpus An open-access, NLP-ready corpus of Catholic texts spanning 2,000 years of the Catholic intellectual tradition — from the Church Fathers to the 20th century. 67,772 content files | 16 collections | 35.9 GB | 2,000 years of coverage What's In the Corpus # Collection Content Files Size Format 01 Git Repos (CSEL, Aquinas Opera Omnia, Septuagint, Byzantine Text, eBible) 4,668 3.4 GB TEI XML, TXT 02 Corpus Corporum (Patrologia Latina + 29… See the full description on the dataset page: https://huggingface.co/datasets/CatholicCorpus/catholiccorpus.
No dataset card provided yet.
Mirrored from an external registry.
Last synced 6/14/2026
Preview not yet available for this dataset