Dataset Card for STSB The Semantic Textual Similarity Benchmark (Cer et al., 2017) is a collection of sentence pairs drawn from news headlines, video and image captions, and natural language inference data. Each pair is human annotated with a similarity score from 1 to 5. However, for this variant, the similarity scores are normalized to between 0 and 1. Dataset Details Columns: "sentence1", "sentence2", "score" Column types: str , str , float Examples: Collection strategy: Reading the sentences and score from STSB dataset and dividing the score by 5. Deduplified: No
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy