GlotLID Description GlotLID is a Fasttext language identification (LID) model that supports more than 2000 labels . Latest: GlotLID is now updated to V3 . V3 supports 2102 labels (three letter ISO codes with script). For more details on the supported languages and performance, as well as significant changes from previous versions, please refer to https://github.com/cisnlp/GlotLID/blob/main/languages v3.md. Demo: huggingface Repository: github Paper: paper (EMNLP 2023) Point of Contact: amir@cis.lmu.de How to use Here is how to use this model to detect the language of a given text: If you are not a fan of huggingface hub, then download the model directyly: License The model is distributed under the Apache License, Version 2.0 plus notices (see LICENSE file for full terms). Version We always maintain the previous version of GlotLID in our repository. To access a specific version, simply append the version number to the filename . For v1: model v1.bin (introduced in the GlotLID paper and used in all experiments). For v2: model v2.bin (an edited version of v1, featuring more languages, and cleaned from noisy corpora based on the analysis of v1). For v3: model v3.bin (an edited version…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy