BACC-18

Bengali Authorship Classification Corpus-18

The files associated with this dataset are licensed under a Creative Commons Attribution 4.0 International license.Introduced 2021-07-09

The developed BACC-18 contains the text of 18 famous authors of Bengali literature. To build this corpus, we crawled texts from four online sources namely NLTR society for natural language technology research [36], Ebanglalibrary [37], Git repository [38] and Blogs [39]–[40][41]. The maximum number of texts (13,308) are collected from NLTR source whereas minimum number of texts (240) are crawled from Blogs. A self-built automatic web crawler3 is used to scrapping the data from four sources. Due to HTML page structure variation of sources, we used various web crawler instead of a typical crawler. In particular, the proposed research has developed 31 Python crawler which can automatically crawl textual data based on the robots.txt policy. The robots.txt policy ensures the search engine whether a crawler can or cannot crawl the particular text contents from a source.4 Initially, we manually selected the famous and authentic web portal’s hyperlink to collect the author’s texts. Web crawler starts with the hyperlink and a spider explore all the pages under the hyperlink to scrapping the author text. After collecting all the authors’ text and we prepared the authorship classification corpus with annotation based on the hyperlink. A single hyperlink contains only a single author text. This hyperlink based web crawling reduces the manual annotation time and cost of human efforts. Corpus Link: https://data.mendeley.com/datasets/y64fcp2nzz