Model Card: Vietnamese Embedding v2 Vietnamese Embedding v2 is an embedding model fine tuned from the BGE M3 model (https://huggingface.co/BAAI/bge reranker v2 m3) to enhance retrieval capabilities for Vietnamese. The model was trained on approximately 1,100,000 triplets of queries, positive documents, and negative documents for Vietnamese. The model was trained with a maximum sequence length of 2304 (256 for query and 2048 for passages). Model Details Model Description Model Type: Sentence Transformer Base model: BAAI/bge m3 Maximum Sequence Length: 2048 tokens Output Dimensionality: 1024 dimensions Similarity Function: Dot product Similarity Language: Vietnamese Licence: Apache 2.0 Usage Evaluation: Dataset: Entire training dataset of Legal Zalo 2021. Our model was not trained on this dataset. Model Accuracy@1 Accuracy@3 Accuracy@5 Accuracy@10 MRR@10 Vietnamese Reranker 0.7944 0.9324 0.9537 0.9740 0.8672 Vietnamese Embedding v2 0.7262 0.8927 0.9268 0.9578 0.8149 Vietnamese Embedding 0.7274 0.8992 0.9305 0.9568 0.8181 Vietnamese bi encoder (BKAI) 0.7109 0.8680 0.9014 0.9299 0.7951 BGE M3 0.5682 0.7728 0.8382 0.8921 0.6822 Vietnamese Reranker and Vietnamese Embedding v2 was trained…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy