license: other license name: idl train license link: LICENSE task categories: image to text size categories: 10M An example page of one pdf document from the Industry Documents Library. This instance of IDL is in webdataset .tar format. Usage with chug Check out chug, our optimized library for sharded dataset loading! Usage with datasets This dataset can also be used with webdataset library or current releases of Hugging Face datasets . Here is an example using the "streaming" parameter. We do recommend downloading the dataset to save bandwidth. For faster download, you can directly use the huggingface hub library. Make sure hf transfer is installed prior to downloading and mind that you have enough space locally. Further, a metadata file pdfa english train info minimal.json contains the list of samples per shard, with same basename and .json or .pdf extension, as well as the count of files per shard. Words and lines document metadata Initially, we obtained the raw data from the IDL API and combined it with the idl data annotation. This information is then reshaped into lines organized in reading order, under the key lines. We keep non reshaped word and bounding box information und…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy