open rdl books Dataset Description Language dan, dansk, Danish License Public Domain, cc0 1.0 Dataset Summary Documents from the Royal Danish Library published between 1750 and 1930. The dataset has each page of each document in image and text format. The text was extracted with OCR. The documents (books of various genres) were obtained from the library. The dataset was assembled to make these public domain Danish texts more accessible. Dataset Structure Data Instances The "page text" was obtained through OCR, and is therefore likely to contain noisy data, especially in older documents, where the original text is either handwritten or printed in fonts, which are not commonly used today, such as Fraktur. "author" and "title" may be missing, especially in documents published before 1833. Those are considered public domain, as there is no probable way the authors could be dead less than 70 years, and no in depth metadata collection was done. "digitalized" may be missing. Data Splits All data is in the "train" split. Data is organized by year of publication, and is segmented into General Data Extraction flowchart is for a broad understanding and is not a fully accurate representation.…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy