PerPaDa is a Persian paraphrase dataset that is collected from users' input in a plagiarism detection system.