Open Library The complete Open Library catalog in clean, analysis ready Parquet. 150.0M+ records across 11 entity types, from ISBNs and author bios to reading logs and Wikidata links. What is it? Open Library is a complete snapshot of the Open Library database, an open project of the Internet Archive with the mission of creating "one web page for every book ever published." The catalog is community edited and contains bibliographic records for millions of authors, works, and physical editions, along with user contributed star ratings and reading logs. This dataset converts the official OpenLibrary Data Dumps from their native TSV+JSON format into clean, columnar Apache Parquet files with Zstd compression. Every field from every record type is fully preserved. Nothing is dropped or filtered. Dump date: 2026 02 Total records: 150.0M License: CC0 1.0 (Public Domain) Why this dataset? Open Library publishes monthly bulk dumps, but they arrive as multi gigabyte gzipped TSV files with embedded JSON. They are awkward to query, impossible to stream into a training pipeline, and painful to join across entity types. This dataset takes care of all that: Columnar : every field is a named Parqu…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy