19,997 machine learning datasets
19,997 dataset results
GamePad that can be used to explore the application of machine learning methods to theorem proving in the Coq proof assistant.
GameWikiSum is a domain-specific (video game) dataset for multi-document summarization, which is one hundred times larger than commonly used datasets, and in another domain than news. Input documents consist of long professional video game reviews as well as references of their gameplay sections in Wikipedia pages.
GASP is a dataset composed by a list of cited abstracts associated with the corresponding source abstract. The dataset is composed by a training set of 100000 elements, a test set and a validation set of 10000 each. The goal is to generate a paper abstract given cited paper's abstracts and model the human creativity behind the process.
A genomics dataset for OOD detection that allows other researchers to benchmark progress on this important problem.
The Gigaword Entailment dataset is a dataset for entailment prediction between an article and its headline. It is built from the Gigaword dataset.
A distant supervision dataset by linking the entire English ClueWeb09 corpus to Freebase.
The Generix Object Zero-shot Learning (GOZ) dataset is a benchmark dataset for zero-shot learning.
A new dataset containing over 550K pairs (covering 143 km^2 area) of RGB and aerial LIDAR depth images.
A large-scale grasp pose detection dataset with a unified evaluation system. The dataset contains 87,040 RGBD images with over 370 million grasp poses.
This is a high-quality dataset consisting of 14.8M utterances in English, extracted from processed dialogues from publicly available online books.
A data set of hourly time phrases from 52,183 fictional books.
HAM is a dataset for molecular graph partitioning. This dataset contains coarse-grained (CG) mappings of 1206 organic molecules with less than 25 heavy atoms. Each molecule was downloaded from the PubChem database as SMILES. One molecule was assigned to two annotators to compare the human agreement between CG mappings. Downloaded SMILES were hand-mapped. The completed annotations were reviewed by a third person, to identify and remove unreasonable mappings (eg: one bead mappings) which did not agree with the given guidelines. Hence, there are 1.68 annotations per molecule in the current database (16% removed).
The HANS (Heuristic Analysis for NLI Systems) dataset which contains many examples where the heuristics fail.
This dataset is used for predicting house prices from both images and textual information. It is composed of 535 sample houses from California, USA.
Ice Hockey News Dataset is a corpus of Finnish ice hockey news, edited to be suitable for training of end-to-end news generation methods, as well as demonstrate generation of text, which was judged by journalists to be relatively close to a viable product.
IgboNLP is a standard machine translation benchmark dataset for Igbo. It consists of 10,000 English-Igbo human-level quality sentence pairs mostly from the news domain.
iLur News Texts is a dataset of over 12000 news articles from iLur.am, categorized into 7 classes: sport, politics, weather, economy, accidents, art, society. The articles are split into train (2242k tokens) and test sets (425k tokens).
Image Caption Quality Dataset is a dataset of crowdsourced ratings for machine-generated image captions. It contains more than 600k ratings of image-caption pairs.
Consists of 10,000 images is constructed, in which all the immediacy measures and the human poses are annotated.
A dataset that consists of 20 actions of various actors, such as tennis serves, yoga and Tai Chi. Take tennis serves as an example. The publicly available videos of some tennis players from YouTube are downloaded, and manually crop the videos roughly to obtain a set of video clips of serves for each player.