CatholicCorpus — Extracted Text Pre extracted plain text from the CatholicCorpus — 2,000 years of the Catholic intellectual tradition, ready for NLP, RAG, and digital humanities. This dataset contains 47,407 plain text files (5.7 GB, 2.64 billion GPT 2 tokens) extracted from the raw source corpus (PDF, EPUB, TEI XML, HTML). If you need the original source formats, see the raw corpus. Quick Start Or clone directly: What's Included Collection Files Description : 01 Git Repos 4,593 Aquinas Opera Omnia, LXX, Byzantine text, eBible 02 Corpus Corporum 9,945 Patrologia Latina, medieval Latin (TEI XML source) 03 Project Gutenberg 8 Catholic classics in English 04 Direct Downloads 1,232 SBL Greek NT, Vulgate, Canon Law, miscellaneous 05 CCEL 54 Ante Nicene and Nicene Fathers 06 Catholic Encyclopedia 15 1913 Catholic Encyclopedia volumes 07 Latin & Franciscan 21 Peter Lombard, Bonaventure, Franciscan authors 08 Liturgical & Hymns 31,090 Divine Office, Gregorian chant texts 09 Archive & Misc 57 Hagiography, scholastics, devotional works 11 Catholic Bible 11 Haydock, Knox, Confraternity editions 12 English Catholic Thinkers 45 Newman, Chesterton, Belloc, Knox 13 Mystics 13 Teresa of Ávila, Joh…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy