SemEval-2015 Task-1

Paraphrase and Semantic Similarity in Twitter

Given two sentences, the participants are asked to determine whether they express the same or very similar meaning and optionally a degree score between 0 and 1. Following the literature on paraphrase identification, we evaluate system performance primarily by the F-1 score and Accuracy against human judgments. We also provide additional evaluations by Pearson correlation and PINC (Chen and Dolan, 2011), which measure lexical dissimilarity between sentence pairs.