Collaborative Open Legal Data (COLD) Cases COLD Cases is a dataset of 8.3 million United States legal decisions with text and metadata, formatted as compressed parquet files. If you'd like to view a sample of the dataset formatted as JSON Lines, you can view one here This dataset exists to support the open legal movement exemplified by projects like Pile of Law and LegalBench. A key input to legal understanding projects is caselaw the published, precedential decisions of judges deciding legal disputes and explaining their reasoning. United States caselaw is collected and published as open data by CourtListener, which maintains scrapers to aggregate data from a wide range of public sources. COLD Cases reformats CourtListener's bulk data so that all of the semantic information about each legal decision (the authors and text of majority and dissenting opinions; head matter; and substantive metadata) is encoded in a single record per decision, with extraneous data removed. Serving in the traditional role of libraries as a standardization steward, the Harvard Library Innovation Lab is maintaining this open source pipeline to consolidate the data engineering for preprocessing caselaw so…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy