NKJP-NER

The NKJP-NER dataset is based on a human-annotated part of the National Corpus of Polish (NKJP). In this dataset, sentences containing named entities of exactly one type have been extracted. The primary task associated with this dataset is to predict the type of the named entity. The dataset provides examples split into three categories:

  1. Train: Contains 15,794 examples.
  2. Validation: Contains 1,941 examples.
  3. Test: Contains 2,058 examples.

The named entity types in this dataset include:

  • geogName
  • noEntity
  • orgName
  • persName
  • placeName
  • time

The NKJP-NER dataset is a valuable resource for natural language processing tasks related to named entity recognition in the Polish language. It is available under the GNU GPL v.3 license and is currently at version 1.1.0¹²³.

Source: Conversation with Bing, 3/16/2024 (1) nkjp-ner | TensorFlow Datasets. https://www.tensorflow.org/datasets/community_catalog/huggingface/nkjp-ner. (2) +86 Ner Datasets - NLP Database - Metatext. https://metatext.io/datasets-list/ner-task. (3) The Best Polish Language Datasets of 2022 | Twine. https://www.twine.net/blog/polish-language-datasets/. (4) nkjp-ner · Datasets at Hugging Face. https://huggingface.co/datasets/nkjp-ner/viewer/default/test. (5) undefined. https://klejbenchmark.com/static/data/klej_nkjp-ner.zip.