Nepali Text Corpus

Introduced 2024-09-14

Overview Nepali-Text-Corpus is a comprehensive collection of approximately 6.4 million articles in the Nepali language. This dataset is the largest text dataset on Nepali Language. It encompasses a diverse range of text types, including news articles, blogs, and more, making it an invaluable resource for researchers, developers, and enthusiasts in the fields of Natural Language Processing (NLP) and computational linguistics.

Dataset Details Total Articles: ~6.4 million Language: Nepali Size: 27.5 GB (in csv) Source: Collected from various Nepali news websites, blogs, and other online platforms.