19,997 machine learning datasets
19,997 dataset results
The National Speech Corpus (NSC) is a significant initiative led by the Info-communications and Media Development Authority (IMDA) of Singapore. It serves as the first large-scale Singapore English corpus and aims to be a valuable resource for researchers and developers working on automatic speech recognition (ASR) technology and other speech-related applications¹²³.
The Safety Prompts dataset is a valuable resource for evaluating and enhancing the safety of large language models (LLMs) in the Chinese language. It consists of carefully crafted prompts that align model outputs with human values, specifically focusing on safety assessment.
Seizures and seizure-like rhythmic and periodic brain activity known as “ictal-interictal-injury continuum” (IIIC) patterns are frequently detected during brain monitoring with electroencephalography (EEG) in patients with epilepsy or critical illness. Prior efforts to automate detection of IIIC patterns have been limited by lack of large well-annotated datasets to train/evaluate algorithms, and there have been only a few attempts to detect IIIC events other than seizures. The IIIC dataset includes 50,697 labeled EEG samples from 2,711 patients’ and 6,095 EEGs that were annotated by physician experts from 18 institutions. These samples were used to train SPaRCNet (Seizures, Periodic and Rhythmic Continuum patterns Deep Neural Network), a computer program that classifies IIIC events with an accuracy matching clinical experts.
UAV tracking dataset with common corruptions tailored for unmannered aerial video
The Benchmark Intended Grouping of Open Speech (BIGOS) is a novel corpus specifically designed for Polish Automatic Speech Recognition (ASR) systems. This initial version of the benchmark comprises 1,900 audio recordings from 71 distinct speakers, sourced from 10 publicly available speech corpora¹²³.
Z-Bench is a fascinating Chinese language model prompt dataset developed by an enthusiastic AI-focused team at Zhenfund. Let me share some intriguing details about it:
This dataset contains 4,403 Indonesian tweets that have been labeled into five emotion classes: love, anger, sadness, joy, and fear. Each line in the dataset consists of a tweet and its respective emotion label separated by a semicolon (,). The first line serves as a header. Pre-processing has been applied to the tweets, including replacing usernames (@) with [USERNAME], URLs/hyperlinks (http://… or https://…) with [URL], and sensitive numbers (e.g., phone numbers, invoice numbers) with [SENSITIVE-NO].
Accompanying software package and data for the publication titled "Bayesian multi-exposure image fusion for robust high dynamic range ptychography".
The increase in religiously motivated hate on social media is clear and ongoing. These platforms have become fertile ground for the dissemination of hate speech directed at religious communities, resulting in tangible repercussions in the real world. Much of the current research concerning the automated identification of hateful content on social media focuses on English-language content. There is comparatively less exploration in low-resource languages such as Hindi. As social media users increasingly utilize their regional languages for expression, it becomes crucial to dedicate appropriate research efforts to hate speech detection in these languages.
Mudestreda Multimodal Device State Recognition Dataset obtained from real industrial milling device with Time Series and Image Data for Classification, Regression, Anomaly Detection, Remaining Useful Life (RUL) estimation, Signal Drift measurement, Zero Shot Flank Took Wear, and Feature Engineering purposes.
Google Research Footbal is a new reinforcement learning environment where agents are trained to play football in an advanced, physics-based 3D simulator. The resulting environment is challenging, easy to use and customize, and it is available under a permissive open-source license. In addition, it provides support for multiplayer and multi-agent experiments.
academy 3 vs 1 with keeper on Google Research Football
YT_subtitles is a remarkable tool designed for building a dataset from YouTube subtitles. Let me break it down for you:
The Hacker News dataset provides a valuable glimpse into the tech industry's landscape. It encompasses posts from Y Combinator's social news website, spanning from 2006 to late 2017¹. Here are some key points about this dataset:
NIH Grant Abstracts: ExPORTER is a valuable resource for researchers and data enthusiasts. It serves as an open data repository containing administrative information about NIH-funded research projects. Here are some key points about ExPORTER:
The Nordjylland News dataset is a collection of news articles from Northern Jutland in Denmark. It provides valuable information about events, developments, and stories specific to that region. Here are some key details about this dataset:
The "no-sammendrag" dataset is a collection of summarized texts in Norwegian. It covers various topics and is useful for tasks related to summarization. The dataset contains documents along with their corresponding summaries. Let me provide more details about it:
The Dutch Social Dataset is a collection of tweets primarily written in Dutch. Let me provide you with some details about this dataset:
The Nordjylland News summarization dataset is a collection of news articles from the Nordjylland region in Denmark, along with their corresponding summaries. This dataset is valuable for training and evaluating text summarization models. Let's delve into the details:
The Swedish ABSAbank-Imm 1.1 is an annotated corpus designed for aspect-based sentiment analysis related to immigration in Sweden. Let's delve into the details: