ANCHOLIK-NER

ANCHOLIK-NER: A Benchmark Dataset for Bangla Regional Named Entity Recognition

CC BY 4.0Introduced 2025-02-16

We developed ANCHOLIK-NER, a Bangla Regional Named Entity Recognition dataset focusing on the Sylhet, Chittagong, and Barishal dialects. It comprises 10,443 sentences, evenly distributed across the three regions, with entities categorized into 10 types. The dataset, sourced from both formal (57.54%) and informal (42.46%) texts, enhances NER models by capturing regional linguistic nuances.