German BERT large paraphrase cosine This is a sentence transformers model. It maps sentences & paragraphs (text) into a 1024 dimensional dense vector space. The model is intended to be used together with SetFit to improve German few shot text classification. It has a sibling model called deutsche telekom/gbert large paraphrase euclidean. This model is based on deepset/gbert large. Many thanks to deepset! Loss Function \ We have used MultipleNegativesRankingLoss with cosine similarity as the loss function. Training Data \ The model is trained on a carefully filtered dataset of deutsche telekom/ger backtrans paraphrase. We deleted the following pairs of sentences: min char len less than 15 jaccard similarity greater than 0.3 de token count greater than 30 en de token count greater than 30 cos sim less than 0.85 Hyperparameters learning rate: 8.345726930229726e 06 num epochs: 7 train batch size: 57 num gpu: 1 Evaluation Results We use the NLU Few shot Benchmark English and German dataset to evaluate this model in a German few shot scenario. Qualitative results multilingual sentence embeddings provide the worst results Electra models also deliver poor results German BERT base size mode…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy