19,997 machine learning datasets
19,997 dataset results
Open KB canonicalization dataset. ReVerb45K increases the entity number to 7.5K and has 45K triples in total. ReVerb45K extract a source sentence for each triple from ClueWeb09
A database of 56 high quality fabric material measurements, provided as carefully calibrated rectified HDR images, together with SVBRDF fits. Used in the Fabric Appearance Challange.
RUHSOLD is hate speech and offensive language dataset in Roman Urdu. The dataset contains over 10 thousand tweets that are hand labelled into the following categories: 1) Abusive/Offensive 2) Untargeted 3) Sexism 4) Religious 5) Neutral
Kinect-WSJ is a multichannel, multispeaker, reverberated, noisy dataset which extends the WSJ0-2mix singlechannel, non-reverberated, noiseless dataset to the strong reverberation and noise conditions and the Kinect-like microphone array geometry used in CHiME-5.
The AbstRCT dataset consists of randomized controlled trials retrieved from the MEDLINE database via PubMed search. The trials are annotated with argument components and argumentative relations.
This dataset contains orthographic samples of words in 19 languages (ar, br, de, en, eno, ent, eo, es, fi, fr, fro, it, ko, nl, pt, ru, sh, tr, zh). Each sample contains two text features: a Word (the textual representation of the word according to its orthography) and a Pronunciation (the highest-surface IPA pronunciation of the word as pronunced in its language).
Maintenance of Wakefulness Test (MWT) is a dataset of recordings with microsleep episodes and drowsiness.
Included in this content:
The Deep Thermal Imaging dataset consists of two main datasets:
Fongbe Data collected by Fréjus A. A LALEYE
The Lens Flare dataset is an internal dataset for Flare Spot detection used in the paper "Automatic Flare Spot Artifact Detection and Removal in Photographs" by Patricia Vitoria and Coloma Ballester.
Sara motion is a 3D motion dataset, named Synthetic Actors and Real Actions (SARA), for training a model to produce motion embeddings suitable for reasoning about motion similarity.
Motion similarity annotations for NTU RGB+D 120 dataset to evaluate motion similarity in the real world.
BU-BIL is an image library which includes six datasets that represent three imaging modalities and six object types. Providers of the datasets are instructed to choose images that capture the various environmental conditions and imaging noise that arose in their studies. These experts are asked to then select objects from those images that reflect the natural diversity of shape and appearances that these objects can exhibit. The image subregions containing the identified objects are cropped to create the image library. The outcome was a library with 305 objects from 235 images. Authors verify by visual inspection that the image library includes a variety of object appearances, backgrounds, and properties distinguishing objects from the background.
Malware Traffic Analysis Knowledge Dataset 2019 (MTA-KDD'19) is an updated and refined dataset specifically tailored to train and evaluate machine learning based malware traffic analysis algorithms. To generate it, that authors started from the largest databases of network traffic captures available online, deriving a dataset with a set of widely-applicable features and then cleaning and preprocessing it to remove noise, handle missing data and keep its size as small as possible. The resulting dataset is not biased by any specific application (although specifically addressed to machine learning algorithms), and the entire process can run automatically to keep it updated.
Data Set Information: The main goal of this data set is providing clean and valid signals for designing cuff-less blood pressure estimation algorithms. The raw electrocardiogram (ECG), photoplethysmograph (PPG), and arterial blood pressure (ABP) signals are originally collected from the physionet.org and then some preprocessing and validation performed on them. (For more information about the process please refer to our paper)
The POTUS Corpus is a Database of Weekly Addresses for the Study of Stance in Politics and Virtual Agents.
The training and validation data are subsets of the training split of the Imagenet 2012. The test set is taken from the validation split of the Imagenet 2012 dataset. Each data set includes 50 images per class.
The National Institute of Informatics provides LIFULL HOME'S Dataset to researchers, which was offered by LIFULL Co., Ltd. for promoting research in informatics and the related fields.