19,997 machine learning datasets
19,997 dataset results
This data contains about 2500 trajectories (with images and actions) of a Sawyer robot interacting with various objects.
MLQuestions is a domain-adaptation dataset for the machine learning domain containing 50K unaligned passages and 35K unaligned questions, and 3K aligned passage and question pairs.
ARC Ukiyo-e Faces is a large-scale (>10k paintings, >20k faces) Ukiyo-e dataset with coherent semantic labels and geometric annotations through augmenting and organizing existing datasets with automatic detection.
This is an complete set of the data we collected and analyzed in our study "Quo Vadis, Open Source? The Limits of Open Source Growth". Please see our GitHub repository for details and tool chain.
Webis-ConcluGen-21 is a large-scale corpus of 136,996 samples of argumentative texts and their conclusions used for the task of generating informative conclusions.
CTFW is a large annotated procedural text dataset in the cybersecurity domain (3154 documents). It is used to generate flow graphs from procedural texts.
The LSEC (Live Stream E-Commerce) dataset has two subsets: LSEC-Small and LSEC-Large. It is a dataset for studying E-commerce transactions in the context of live streams, where the streames are talking about products while interacting with their audience. The dataset consists of interaction information among streamers, users, and products.
Unsplash2K is high-resolution image dataset with 2K resolution. Unsplash2K dataset is crawled from unsplash. Unsplash2K dataset contains 498 high-resolution images and corresponding low-resolution images which are downsampled by bicubic downsamling for x2, x4, x8 scale. Unsplash2K contains diverse contents such as animals, architectures and flowers.
The dataset provides News articles obtained from emol.cl including their content, title and all the comments it received in JSON format
Structure of code/data folders and how to use them fastzip-code
A collection of photographic and synthetic images intended for analysis of image processing techniques and quality assessment of displays.
Number of images: 1,657 images during or after the fire
100,000 LR synthetic barcode datasets along with their corresponding bounding boxes ground truth masks.
~1M Flickr images from the XX century-aged from the 1910s to 1990s. Dataset was introduced by Müller et al. and can be found https://www.radar-service.eu/radar/en/dataset/tJzxrsYUkvPklBOw
This is a pose estimation dataset, consisting of symmetric 3D shapes where multiple orientations are visually indistinguishable. The challenge is to predict all equivalent orientations when only one orientation is paired with each image during training (as is the scenario for most pose estimation datasets). In contrast to most pose estimation datasets, the full set of equivalent orientations is available for evaluation.
Intermediate annotations from the FEVER dataset that describe original facts extracted from Wikipedia and the mutations that were applied, yielding the claims in FEVER.
This dataset contains the bus trajectory dataset collected by 6 volunteers who were asked to travel across the sub-urban city of Durgapur, India, on intra-city buses (route name: 54 Feet). During the travel, the volunteers captured sensor logs through an Android application installed on COTS smartphones.
MARS dataset processed with our re-Detect and Link (DL) module.
DukeMTMC-VideoReID-DL processed with our re-Detect and Link (DL) module.