Model Card for LT MLKM modernBERT Table of Contents Model Details How to Get Started with the Model Uses Risks, Biases, and Limitations Training Details Evaluation Citation License Model Details Model name: LT MLKM modernBERT Project: “Development of the General Lithuanian Language Corpus and Vectorized Lithuanian Language Models” carried out by the State Digital Solutions Agency (SDSA) (Contract No. VDU S 1684). The SDSA project manager is A. Rakauskas, and the supplier group leader is Assoc. Prof. Dr. A. Utka. Architecture: ModernBERT from NVIDIA. Model description: LT MLKM modernBERT is a Lithuanian masked language model developed as part of the national project “Development of a general Lithuanian language corpus and vectorized models.” The model builds on the ModernBERT base architecture and was pre trained on the BLKT Lithuanian Text Corpus Stage 3, which includes over 1.87 billion words and 49 billion training tokens from diverse Lithuanian sources such as news, legal, academic, and public sector texts. With a context length of 8,192 tokens, it efficiently processes long documents while maintaining linguistic precision and coherence. The model advances the project’s goal of…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy