SweDN
The SweDN 1.0 dataset is a valuable resource for natural language processing (NLP) tasks, specifically text summarization. Let's delve into the details:
-
Title and Subtitle:
- Title: SweDN 1.0
- Subtitle: A Swedish text summarization corpus
-
Description:
- The SweDN 1.0 corpus is based on 1,963,576 news articles from the Swedish newspaper Dagens Nyheter (DN) spanning the years 2000 to 2020.
- These articles have been filtered to resemble the CNN/DailyMail dataset in terms of their textual structure.
-
Purpose and Usage:
- Model Development: SweDN 1.0 serves as a training resource for both extractive and abstractive text summarizers.
- Intended Task: Given a text (article), the goal is to provide its summary.
- Evaluation Measures: The recommended evaluation metrics include the harmonic mean of Bleu and Rouge, along with Rouge, BERTScore, and Coh-Metrix.
-
Data Details:
- Language: Swedish
- Number of Articles: The dataset comprises 38,121 news articles along with their corresponding preambles.
- Format: The data is available in JSONL and TSV files, containing fields such as ID, headline, summary, article, and article category.
- An additional file provides various statistics for each entry, including length measures, embedding similarity, and article category.
-
Ethical Considerations:
- The dataset does not involve any specific data labeling or annotator characteristics.
- As with any NLP dataset, it's essential to consider ethical aspects and potential biases.
-
References:
- Monsen, J., & Jönsson, A. (2021). A method for building non-English corpora for abstractive text summarization. Proceedings of the CLARIN Annual Conference ¹.
(1) SweDN 1.0 | Språkbanken Text - Göteborgs universitet. https://spraakbanken.gu.se/en/resources/swedn. (2) Language resources | Språkbanken Text - Göteborgs universitet. https://spraakbanken.gu.se/en/resources/train. (3) Kylberg Texture Dataset v. 1.0 – Kylberg.org. https://kylberg.org/kylberg-texture-dataset-v-1-0/. (4) undefined. https://spraakbanken.gu.se/resurser/superlim.