LaBSE Model description Smaller Language agnostic BERT Sentence Encoder (LaBSE) is a BERT based model distilled from the original LaBSE model to 15 languages (from the original 109 languages) using the techniques described in the paper 'Load What You Need: Smaller Versions of Multilingual BERT' by Ukjae Jeong. Model: HuggingFace's model hub. Original model: TensorFlow Hub. Distillation source: GitHub. Conversion from TensorFlow to PyTorch: GitHub. Usage Using the model: To get the sentence embeddings, use the pooler output: Output for other languages: For similarity between sentences, an L2 norm is recommended before calculating the similarity: Details Details about data, training, evaluation and performance metrics are available in the original paper. BibTeX entry and citation info
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy