Model Card for Indus Retriever Indus Retriever ( nasa smd ibm st v2 ) is a Bi encoder sentence transformer model, that is fine tuned from nasa smd ibm v0.1 encoder model. it is an updated version of nasa smd ibm st with better performance (shown below). It's trained with 271 million examples along with a domain specific dataset of 2.6 million examples from documents curated by NASA Science Mission Directorate (SMD). With this model, we aim to enhance natural language technologies like information retrieval and intelligent search as it applies to SMD NLP applications. you can also use distilled version of the model here: https://huggingface.co/nasa impact/nasa ibm st.38m Model Details Base Encoder Model : INDUS Tokenizer : Custom Parameters : 125M Training Strategy : Sentence Pairs, and score indicating relevancy. The model encodes the two sentence pairs independently and cosine similarity is calculated. the similarity is optimized using the relevance score. Training Data Figure: Open dataset sources for sentence transformers (269M in total) Additionally, 2.6M abstract + title pairs collected from NASA SMD documents. Training Procedure Framework : PyTorch 1.9.1 sentence transformers…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy