19,997 machine learning datasets
19,997 dataset results
A dataset for studying situated goal-directed human communication.
The PREdiction of Clinical Outcomes from Genomic profiles (or PRECOG) encompasses 166 cancer expression data sets, including overall survival data for ~18,000 patients diagnosed with 39 distinct malignancies.
A synthetic dataset with 206K pressure images with 3D human poses and shapes.
Procedural Human Action Videos contains a total of 39,982 videos, with more than 1,000 examples for each action of 35 categories.
A novel stance detection dataset covering 419 different controversial issues and their related pros and cons collected by procon.org in nonpartisan format.
Dataset that can be used to evaluate both general semantic flow techniques and region-based approaches such as proposal flow.
This is a large-scale court judgment dataset, where each judgment is a summary of the case description with a patternized style. It contains 2,003,390 court judgment documents. The case description is used as the input, and the court judgment as the summary. The average lengths of the input documents and summaries are 595.15 words and 273.57 words respectively.
The public_meetings corpus contains meetings, made of pairs of automatic transcriptions from audio recordings and meeting reports written by a professional. 22 aligned meetings are provided in total.
A high-quality large-scale dataset consisting of 49,000+ data samples for the task of Chinese query-based document summarization.
RadioTalk is a corpus of speech recognition transcripts sampled from talk radio broadcasts in the United States between October of 2018 and March of 2019. The corpus is intended for use by researchers in the fields of natural language processing, conversational analysis, and the social sciences. The corpus encompasses approximately 2.8 billion words of automatically transcribed speech from 284,000 hours of radio, together with metadata about the speech, such as geographical location, speaker turn boundaries, gender, and radio program information.
Corpus for improving the quality and relational abilities of Intelligent Virtual Agents (IVAs).
A dataset to encourage the community to adapt oriented bounding box (OBB) detectors for more complex environments.
This dataset is used for RF signal recognition, used to recognize different RF devices based on the signals they transmitted.
A benchmark suite of continuous control tasks, including classic tasks like cart-pole swing-up, tasks with very high state and action dimensionality such as 3D humanoid locomotion, tasks with partial observations, and tasks with hierarchical structure.
A large dataset from games of some of the top teams (from 2016 and 2017) in RoboCup Soccer Simulation League (2D), where teams of 11 robots (agents) compete against each other.
Tagged for Sentiment (Positive, Negative, Neutral).
The RotoWire-Modified dataset is a cleaned extension of the RotoWire dataset, with writer information about each document. It contains 2705 samples for training, 532 for validation and 497 for testing.
A corpus of real-world spoken personal narratives comprising 10,296 narrative clauses from 594 video transcripts.
The data consists of 100,000 renderings each of the Bunny and Dragon objects from the Stanford 3D Scanning Repository. More objects may be added in the future, but only the Bunny and Dragon are used in the paper. Each object is rendered with a uniformly sampled illumination from a point on the 2-sphere, and a uniformly sampled 3D rotation. The true latent states are provided as NumPy arrays along with the images. The lighting is given as a 3-vector with unit norm, while the rotation is provided both as a quaternion and a 3x3 orthogonal matrix.
A dataset for grounded language learning that consists of navigational instructions and actions in a maze-like environment.