Model Description: vietnamese embedding is the Embedding Model for Vietnamese language. This model is a specialized sentence embedding trained specifically for the Vietnamese language, leveraging the robust capabilities of PhoBERT, a pre trained language model based on the RoBERTa architecture. The model utilizes PhoBERT to encode Vietnamese sentences into a 768 dimensional vector space, facilitating a wide range of applications from semantic search to text clustering. The embeddings capture the nuanced meanings of Vietnamese sentences, reflecting both the lexical and contextual layers of the language. Full Model Architecture Training and Fine tuning process The model underwent a rigorous four stage training and fine tuning process, each tailored to enhance its ability to generate precise and contextually relevant sentence embeddings for the Vietnamese language. Below is an outline of these stages: Stage 1: Initial Training Dataset: ViNLI SimCSE supervised Method: Trained using the SimCSE approach which employs a supervised contrastive learning framework. The model was optimized using Triplet Loss to effectively learn from high quality annotated sentence pairs. Stage 2: Continued F…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy