NSC

National Speech Corpus

The National Speech Corpus (NSC) is a significant initiative led by the Info-communications and Media Development Authority (IMDA) of Singapore. It serves as the first large-scale Singapore English corpus and aims to be a valuable resource for researchers and developers working on automatic speech recognition (ASR) technology and other speech-related applications¹²³.

Here are some key points about the National Speech Corpus:

  1. Purpose: The NSC is designed to become an essential source of open speech data for ASR research and speech-related applications.
  2. Importance: By providing a large-scale corpus of Singapore English, it enables advancements in speech recognition, speech synthesis, and other related technologies.
  3. AI and Digital Solutions: The NSC harnesses the power of Artificial Intelligence (AI), paving the way for innovative digital solutions and driving progress in Singapore's digital landscape.
  4. Accent Recognition: Supporting speech technologies can sometimes struggle with recognizing and transcribing locally accented English. The NSC addresses this gap by improving speech engines' accuracy for locally accented English.
  5. Speech Synthesis: The NSC also contributes to speech synthesis technology, allowing AI voices to be produced with more accurate local pronunciations and familiarity for Singaporeans.
  6. Applications: As speech technology improves, ASR can be used in various applications, such as telco call centers, chatbots, and more. For instance:
    • Telco call centers can use ASR to transcribe calls for auditing and sentiment analysis purposes.
    • Chatbots can go beyond text and accurately support our accent while replying in a familiar local accent with accurate pronunciations of street names and food.

The NSC is available under the Singapore Open Data License, and researchers interested in using it can download the corpus via a Dropbox account. The current size of the NSC is approximately 1.2 TB, and it continues to contribute to advancements in speech technology¹³.

Source: Conversation with Bing, 3/17/2024 (1) National Speech Corpus - Infocomm Media Development Authority. https://www.imda.gov.sg/how-we-can-help/national-speech-corpus. (2) ISCA - ArchiveRedirection - isca-speech.org. https://www.isca-speech.org/archive/interspeech_2019/koh19_interspeech.html. (3) National Speech Corpus - Infocomm Media Development Authority. https://www.imda.gov.sg/about-imda/emerging-technologies-and-research/artificial-intelligence/national-speech-corpus.