ModernBERT Ja 130M This repository provides Japanese ModernBERT trained by SB Intuitions. ModernBERT is a new variant of the BERT model that combines local and global attention, allowing it to handle long sequences while maintaining high computational efficiency. It also incorporates modern architectural improvements, such as RoPE. Our ModernBERT Ja 130M is trained on a high quality corpus of Japanese and English text comprising 4.39T tokens , featuring a vocabulary size of 102,400 and a sequence length of 8,192 tokens. How to Use You can use our models directly with the transformers library v4.48.0 or higher: Additionally, if your GPUs support Flash Attention 2, we recommend using our models with Flash Attention 2. Example Usage Model Series We provide ModernBERT Ja in several model sizes. Below is a summary of each model. ID Param. Param. w/o Emb. Dim. Inter. Dim. Layers sbintuitions/modernbert ja 30m 37M 10M 256 1024 10 sbintuitions/modernbert ja 70m 70M 31M 384 1536 13 sbintuitions/modernbert ja 130m 132M 80M 512 2048 19 sbintuitions/modernbert ja 310m 315M 236M 768 3072 25 For all models, the vocabulary size is 102,400, the head dimension is 64, and the activation function is…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy