DEk21 hcmute embedding DEk21 hcmute embedding is a Vietnamese text embedding focused on RAG and production efficiency: 📚 Trained Dataset : The model was trained on an in house dataset consisting of approximately 100,000 examples of legal questions and their related contexts. ⚙️ Efficiency: Trained with a Matryoshka loss , allowing embeddings to be truncated with minimal performance loss. This ensures that smaller embeddings are faster to compare, making the model efficient for real world production use. Model Details Model Description Model Type: Sentence Transformer Maximum Sequence Length: 256 tokens Output Dimensionality: 768 dimensions Similarity Function: Cosine Similarity Language: vietnamese License: apache 2.0 Model Sources Documentation: Sentence Transformers Documentation Repository: Sentence Transformers on GitHub Hugging Face: Sentence Transformers on Hugging Face Full Model Architecture Usage Direct Usage (Sentence Transformers) First install the Sentence Transformers library: Then you can load this model and run inference. Evaluation Metrics Information Retrieval Datasets: another symato/VMTEB Zalo legel retrieval wseg Evaluated with InformationRetrievalEvaluator mod…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy