19,997 machine learning datasets
19,997 dataset results
There is currently much interest and activity aimed at building powerful multi-purpose information systems. The agencies involved include DARPA, ARDA and NIST. Their programmes, for example DARPA's TIDES (Translingual Information Detection Extraction and Summarization) programme, ARDA's Advanced Question & Answering Program and NIST's TREC (Text Retrieval Conferences) programme cover a range of subprogrammes. These focus on different tasks requiring their own evaluation designs.
Data was collected from Tobii Fusion screen-based Eye Tracker. This study collected drivers’ gaze data by letting participants watch dashcam captured videos of driving scenes in the lab setting. Original vidoes are downloaded from https://github.com/Cogito2012/CarCrashDataset. Each video lasts 5 seconds and the frequency of the videos is 10 Hz.
The Euro-PVI dataset contains trajectories of pedestrians and bicyclists, with dense interactions with the ego-vehicle. The dataset is collected in Brussels and Leuven, Belgium. The goal of this dataset is to address the challenge of future trajectory prediction in urban environments with dense pedestrian (bicyclist) - vehicle interactions.
Knot128 is a dataset to test knot untangling algorithms, i.e., highly-tangled configurations that can be difficult to smooth out into a canonical knot embedding. Knot128 is comprised of knots from 128 different isotopy classes; for each class, a tangled embedding, and a canonical embedding are provided.
Trefoil100 is a dataset to test knot untangling algorithms, i.e., highly-tangled configurations that can be difficult to smooth out into a canonical knot embedding. Trefoil100 contains 100 tangled embeddings of the trefoil knot.
This repo contains open-source channel measurement data for research and development purposes.
Introduced by Singh, Sumeet S.. “Teaching Machines to Code: Neural Markup Generation with Visual Attention.” ArXiv abs/1802.05415 (2018): n. pag.
Introduced by Singh, Sumeet S.. “Teaching Machines to Code: Neural Markup Generation with Visual Attention.” ArXiv abs/1802.05415 (2018): n. pag.
The dataset contains summary statistics and engagement metrics captured from users in a live, 'in-the-wild' study of an interactive TV show.
Multispectral and HD vineyard orthomosaics from central Portugal
We provide video sequences with annotated object masks for video inpainting. The resolution is 3840 x 2160.
The MedLEA package provides morphological and structural features of 471 medicinal plant leaves and 1099 leaf images of 31 species and 29-45 images per species.
A dataset of images containing leaves from 15 tree classes.
The dataset is a private dataset collected for automatic analysis of psychological distress. It contains self-reported distress labels provided by human volunteers. The dataset consists of 30-min interview recordings of participants.
XA Bin-Picking is a point-cloud dataset comprising both simulated and real-world scenes with three industrial parts. Synthesized scenes consists of 1000 training samples. The test samples are real scenes and the ground truth instance labels are made manually. There are 20 to 30 identical types of parts randomly piled up in a scene. Each scene contains about 60,000 boundary points. Each point in the scene has instance annotations. The parts are texture-less and have no discernible color. Both of training samples and test sam- ples only contain the boundary points of parts.
This dataset is used for neural co-training. mtl_monkey_dataset: was used for our MTL-Monkey model and involves neural responses that were predicted by a single-task trained model on real monkey V1. mtl_oracle_dataset: was used for our MTL-Oracle model and involves neural responses that were predicted by our image classification oracle. mtl_shuffled_dataset: was used for our MTL-Shuffled model and is the result of shuffling the mtl_monkey_dataset across images.
Bosch Industrial Depth Completion Dataset (BIDCD) is an RGBD dataset for of static table-top scenes with industrial objects. The data was collected with a RealSense depth-camera mounted on a robotic arm, i.e. from multiple Points-of-View (POV), approximately 60 for each scene. We generated depth ground truth with a customized pipeline for removing erroneous depth values, and applied Multi-View geometry to fuse the cleaned depth frames and fill-in missing information. The fused scene mesh was back-projected to each POV, and finally a bi-lateral filter was applied to reduce the remaining holes.
A Bambara dialectal dataset dedicated for Sentiment Analysis, available freely for Natural Language Processing research purposes
Trope Understanding in Movies and Animations (TrUMAn) is a dataset intending to evaluate and develop learning systems beyond visual signals.
A high-resolution multi-sensor remote sensing scene classification dataset, appropriate for training and evaluating image classification models in the remote sensing domain.