19,997 machine learning datasets
19,997 dataset results
This dataset contains around 218K sentences, with 1.5 million words, from 30 different books designed for Post-OCR text correction.
INDRA is a dataset capturing videos of Indian roads from the pedestrian point-of-view. INDRA contains 104 videos comprising of 26k 1080p frames, each annotated with a binary road crossing safety label and vehicle bounding boxes.
MM-Locate-News is a dataset for location estimation of news. It consists of 6395 news articles covering 237 cities and 152 countries across all continents as well as multiple domains such as health, environment, and politics. The dataset is collected in a weakly-supervised manner, and multiple data cleaning steps are applied to remove articles with potential inaccurate geolocation information. The acquired dataset addresses drawbacks of other datasets such as BreakingNews as it considers multimodal content of news to label the corresponding location.
Nlakh is a dataset for Musical Instrument Retrieval. It is a combination of the NSynth dataset, which provides a large number of instruments, and the Lakh dataset, which provides multi-track MIDI data.
YM2413-MDB is an 80s FM video game music dataset with multi-label emotion annotations. It includes 669 audio and MIDI files of music from Sega and MSX PC games in the 80s using YM2413, a programmable sound generator based on FM. The collected game music is arranged with a subset of 15 monophonic instruments and one drum instrument.
PaintNet is a dataset for learning robotic spray painting of free-form 3D objects. PaintNet includes more than 800 object meshes and the associated painting strokes collected in a real industrial setting.
ec-darkpattern is a dataset for dark pattern detection and prepared its baseline detection performance with state-of-the-art machine learning methods. The original dataset was obtained from Mathur et al.’s study in 2019 [11kScale], which consists of 1,818 dark pattern texts from shopping sites. Negative samples, i.e., non-dark pattern texts, by retrieving texts from the same websites as Mathur et al.'s dataset.
It is a competition on kaggle with stroke Prediction, which is heavily imbalanced.
Virtuoso Strings is a dataset for soft onsets detection for string instruments. It consists of over 144 recordings of professional performances of an excerpt from Haydn's string quartet Op. 74 No. 1 Finale, each with corresponding individual instrumental onset annotations.
The main goal of the Continuous Integration of Performance Models (CIPM) is to enable an accurate architecture-based performance prediction at each point of the systems development life cycle. For this goal, the CIPM approach continuously updates the architecture-level performance models of a software system according to observed changes at development time and at operation time.
Comet is a dataset which contains 11.5k user-assistant dialogs (totalling 103k utterances), grounded in simulated personal memory graphs.
Stack of 2D gray images of glass fiber-reinforced polyamide 66 (GF-PA66) 3D X-ray Computed Tomography (XCT) specimen.
The Vesta dataset was released for use in the IEEE CIS Fraud Detection competition. It contains 590,540 card transactions, 20,663 of which are fraudulent (3.5%). Each transaction has 431 features (400 numerical, 31 categorical), along with the relative timestamp and a label of whether it was fraudulent or legitimate. For anonymization purposes, the names of the identity features have been masked, along with the names of the extra features engineered by Vesta.
This dataset contains the main data set of our FOCI 2020 paper "Padding Ain’t Enough: Assessing the Privacy Guarantees of Encrypted DNS".
This dataset contains one part for the "Subpage-Agnostic Domain Classification" section of our FOCI 2020 paper "Padding Ain’t Enough: Assessing the Privacy Guarantees of Encrypted DNS".
This dataset contains the second part of the "Subpage-Agnostic Domain Classification" section of our FOCI 2020 paper "Padding Ain’t Enough: Assessing the Privacy Guarantees of Encrypted DNS".
This dataset contains recordings of 32 sound producing insect species with a total 335 files and a length of 57 minutes. The dataset was compiled for training neural networks to automatically identify insect species while comparing adaptive, waveform-based frontends to conventional mel-spectrogram frontends for audio feature extraction. This work will be submitted for publication in the future and this dataset can be used to replicate the results, as well as other uses. The scripts for audio processing and the machine learning implementations will be published on Github.
SummZoo, a benchmark consists of 8 diverse summarization tasks with multiple sets of few-shot samples for each task, covering both monologue and dialogue domains.
IDK-MRC is an Indonesian Machine Reading Comprehension (MRC) dataset consists of more than 10K questions in total with over 5K unanswerable questions with diverse question types.
Chinese Character Stroke Extraction (CCSE) is a benchmark containing two large-scale datasets: Kaiti CCSE (CCSE-Kai) and Handwritten CCSE (CCSE-HW). It is designed for stroke extraction problems.