Model Description: vietnamese document embedding is the Document Embedding Model for Vietnamese language with context length up to 8096 tokens. This model is a specialized long text embedding trained specifically for the Vietnamese language, which is built upon gte multilingual and trained using the Multi Negative Ranking Loss, Matryoshka2dLoss and SimilarityLoss. Full Model Architecture Training and Fine tuning process The model underwent a rigorous four stage training and fine tuning process, each tailored to enhance its ability to generate precise and contextually relevant sentence embeddings for the Vietnamese language. Below is an outline of these stages: Stage 1: Training NLI on dataset XNLI: Dataset: XNLI vn Method: Training using Multi Negative Ranking Loss and Matryoshka2dLoss. This stage focused on improving the model's ability to discern and rank nuanced differences in sentence semantics. Stage 2: Fine tuning for Semantic Textual Similarity on STS Benchmark Dataset: STSB vn Method: Fine tuning specifically for the semantic textual similarity benchmark using Siamese BERT Networks configured with the 'sentence transformers' library. This stage honed the model's precision i…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy