Dataset Card for The Harvard USPTO Patent Dataset (HUPD) Dataset Description Homepage: https://patentdataset.org/ Repository: HUPD GitHub repository Paper: HUPD arXiv Submission Point of Contact: Mirac Suzgun Dataset Summary The Harvard USPTO Dataset (HUPD) is a large scale, well structured, and multi purpose corpus of English language utility patent applications filed to the United States Patent and Trademark Office (USPTO) between January 2004 and December 2018. Experiments and Tasks Considered in the Paper Patent Acceptance Prediction : Given a section of a patent application (in particular, the abstract, claims, or description), predict whether the application will be accepted by the USPTO. Automated Subject (IPC/CPC) Classification : Predict the primary IPC or CPC code of a patent application given (some subset of) the text of the application. Language Modeling : Masked/autoregressive language modeling on the claims and description sections of patent applications. Abstractive Summarization : Given the claims or claims section of a patent application, generate the abstract. Languages The dataset contains English text only. Domain Patents (intellectual property). Dataset Curator…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy