Model Card: Vietnamese Embedding Vietnamese Embedding is an embedding model fine tuned from the BGE M3 model (https://huggingface.co/BAAI/bge m3) to enhance retrieval capabilities for Vietnamese. The model was trained on approximately 300,000 triplets of queries, positive documents, and negative documents for Vietnamese. The model was trained with a maximum sequence length of 2048. Model Details Model Description Model Type: Sentence Transformer Base model: BAAI/bge m3 Maximum Sequence Length: 2048 tokens Output Dimensionality: 1024 dimensions Similarity Function: Dot product Similarity Language: Vietnamese Licence: Apache 2.0 Usage Evaluation: Dataset: Entire training dataset of Legal Zalo 2021. Our model was not trained on this dataset. Model Accuracy@1 Accuracy@3 Accuracy@5 Accuracy@10 MRR@10 Vietnamese Reranker 0.7944 0.9324 0.9537 0.9740 0.8672 Vietnamese Embedding v2 0.7262 0.8927 0.9268 0.9578 0.8149 Vietnamese Embedding (public) 0.7274 0.8992 0.9305 0.9568 0.8181 Vietnamese bi encoder (BKAI) 0.7109 0.8680 0.9014 0.9299 0.7951 BGE M3 0.5682 0.7728 0.8382 0.8921 0.6822 Vietnamese Reranker and Vietnamese Embedding v2 was trained on 1100000 triplets. Although the score on the l…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy