The text embedding set trained by Jina AI . Quick Start The easiest way to starting using jina embeddings v2 base code is to use Jina AI's Embedding API. Intended Usage & Model Info jina embeddings v2 base code is an multilingual embedding model speaks English and 30 widely used programming languages . Same as other jina embeddings v2 series, it supports 8192 sequence length. jina embeddings v2 base code is based on a Bert architecture (JinaBert) that supports the symmetric bidirectional variant of ALiBi to allow longer sequence length. The backbone jina bert v2 base code is pretrained on the github code dataset. The model is further trained on Jina AI's collection of more than 150 millions of coding question answer and docstring source code pairs. These pairs were obtained from various domains and were carefully selected through a thorough cleaning process. The embedding model was trained using 512 sequence length, but extrapolates to 8k sequence length (or even longer) thanks to ALiBi. This makes our model useful for a range of use cases, especially when processing long documents is needed, including technical question answering and code search. This model has 161 million paramet…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy