MINDS-14

MINDS-14 is a dataset designed for the intent detection task with spoken data. It encompasses 14 distinct intents extracted from a commercial system in the e-banking domain. These intents are associated with spoken examples in 14 diverse language varieties. The dataset serves as a valuable resource for training and evaluating intent detection models.

Here are some key details about the MINDS-14 dataset:

  • Tasks: The primary task is Automatic Speech Recognition, specifically keyword-spotting.
  • Languages: The dataset includes examples in English, French, Italian, and nine other languages.
  • Multilinguality: It is a multilingual dataset.
  • Size Categories: The dataset size falls within the range of 10,000 to 100,000 examples.
  • Language Creators: The annotations were created by a combination of crowdsourced and expert-generated contributors.
  • License: The dataset is available under the CC-BY-4.0 license.

Source: Conversation with Bing, 3/18/2024 (1) PolyAI/minds14 · Datasets at Hugging Face. https://huggingface.co/datasets/PolyAI/minds14. (2) PolyAI/minds14 at main - Hugging Face. https://huggingface.co/datasets/PolyAI/minds14/tree/main. (3) minds14.py · PolyAI/minds14 at main - Hugging Face. https://huggingface.co/datasets/PolyAI/minds14/blob/main/minds14.py.