🧨 Danish Dynaword Version 1.2.16 (Changelog) Language dan, dansk, Danish License Openly Licensed, See the respective dataset Models For model trained used this data see danish foundation models Contact If you have question about this project please create an issue here Table of Contents 🧨 Danish Dynaword Table of Contents Dataset Description Dataset Summary Loading the dataset Languages Domains Licensing Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Annotations Source Data Data Collection and Processing Dataset Statistics Contributing to the dataset Citation Information License information Personal and Sensitive Information Bias, Risks, and Limitations Notice and takedown policy Dataset Description Number of samples : 5.66M Number of tokens (Llama 3) : 6.83B Average document length in tokens (min, max) : 1.21K (2, 12.20M) Dataset Summary The Danish dynaword is a collection of Danish free form text datasets from various domains. All of the datasets in Danish Dynaword are openly licensed and deemed permissible for training large language models. Danish Dynaword is continually developed, which means that the dataset will actively be upd…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy