The ArGiMI Ardian datasets : Text only version The ArGiMi project is committed to open source principles and data sharing. Thanks to our generous partners, we are releasing several valuable datasets to the public. Dataset description This text only dataset comprises 34,000 financial annual reports, written in English, meticulously extracted from their original PDF format to provide a valuable resource for researchers and developers in financial analysis and natural language processing (NLP). These reports were published from the late 90s to 2023. This is the lightweight version without page images. For the full version with page screenshots, see artefactory/Argimi Ardian Finance 10k text image . You can load the dataset with: Dataset composition: Each PDF was divided into individual pages to facilitate granular analysis. For each page, the following data points were extracted: Raw Text: The complete textual content of the page, capturing all textual information present. Cells: Each cell within tables was identified and represented as a Cell object within the docling framework. Each Cell object encapsulates: id : A unique identifier assigned to each cell, ensuring unambiguous refere…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy