SweDN

The SweDN 1.0 dataset is a valuable resource for natural language processing (NLP) tasks, specifically text summarization. Let's delve into the details:

  1. Title and Subtitle:

    • Title: SweDN 1.0
    • Subtitle: A Swedish text summarization corpus
  2. Description:

    • The SweDN 1.0 corpus is based on 1,963,576 news articles from the Swedish newspaper Dagens Nyheter (DN) spanning the years 2000 to 2020.
    • These articles have been filtered to resemble the CNN/DailyMail dataset in terms of their textual structure.
  3. Purpose and Usage:

    • Model Development: SweDN 1.0 serves as a training resource for both extractive and abstractive text summarizers.
    • Intended Task: Given a text (article), the goal is to provide its summary.
    • Evaluation Measures: The recommended evaluation metrics include the harmonic mean of Bleu and Rouge, along with Rouge, BERTScore, and Coh-Metrix.
  4. Data Details:

    • Language: Swedish
    • Number of Articles: The dataset comprises 38,121 news articles along with their corresponding preambles.
    • Format: The data is available in JSONL and TSV files, containing fields such as ID, headline, summary, article, and article category.
    • An additional file provides various statistics for each entry, including length measures, embedding similarity, and article category.
  5. Ethical Considerations:

    • The dataset does not involve any specific data labeling or annotator characteristics.
    • As with any NLP dataset, it's essential to consider ethical aspects and potential biases.
  6. References:

    • Monsen, J., & Jönsson, A. (2021). A method for building non-English corpora for abstractive text summarization. Proceedings of the CLARIN Annual Conference ¹.

(1) SweDN 1.0 | Språkbanken Text - Göteborgs universitet. https://spraakbanken.gu.se/en/resources/swedn. (2) Language resources | Språkbanken Text - Göteborgs universitet. https://spraakbanken.gu.se/en/resources/train. (3) Kylberg Texture Dataset v. 1.0 – Kylberg.org. https://kylberg.org/kylberg-texture-dataset-v-1-0/. (4) undefined. https://spraakbanken.gu.se/resurser/superlim.