Ruri: Japanese General Text Embeddings Ruri v3 is a general purpose Japanese text embedding model built on top of ModernBERT Ja . Ruri v3 offers several key technical advantages: State of the art performance for Japanese text embedding tasks. Supports sequence lengths up to 8192 tokens Previous versions of Ruri (v1, v2) were limited to 512. Expanded vocabulary of 100K tokens , compared to 32K in v1 and v2 The larger vocabulary make input sequences shorter, improving efficiency. Integrated FlashAttention , following ModernBERT's architecture Enables faster inference and fine tuning. Tokenizer based solely on SentencePiece Unlike previous versions, which relied on Japanese specific BERT tokenizers and required pre tokenized input, Ruri v3 performs tokenization with SentencePiece only—no external word segmentation tool is required. Model Series We provide Ruri v3 in several model sizes. Below is a summary of each model. ID Param. Param. w/o Emb. Dim. Layers Avg. JMTEB cl nagoya/ruri v3 30m 37M 10M 256 10 74.51 cl nagoya/ruri v3 70m 70M 31M 384 13 75.48 cl nagoya/ruri v3 130m 132M 80M 512 19 76.55 cl nagoya/ruri v3 310m 315M 236M 768 25 77.24 Usage You can use our models directly with…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy