19,997 machine learning datasets
19,997 dataset results
CMWD (Cloud Motion Wind Dataset) is the first cloud motion wind dataset for deep learning research. It contains 6388 adjacent grayscale image pairs for training and another 715 images pairs for testing.
TCLD (Typhoon Center Location Dataset) is a brand new typhoon center location dataset for deep learning research. It contains 1809 grayscale images for training and another 319 images for testing.
SCMD dataset is a brand new cloudage nowcasting dataset for deep learning research. It contains 20000 grayscale image sequences for training and another 3500 image sequences for testing. You can get the SCMD2016 dataset at any time but only for scientific research. At the same time, please cite our work when you use the SCMD dataset
The dataset contains synthetic training, validation and test data for occupancy grid mapping from lidar point clouds. Additionally, real-world lidar point clouds from a test vehicle with the same lidar setup as the simulated lidar sensor is provided. Point clouds are stored as PCD files and occupancy grid maps are stored as PNG images whereas one image channel describes evidence for a free and another one describes evidence for occupied cell state.
We introduce a first Vietnamese Spelling Correction dataset containing manual labelling mistakes and corresponding correct words.
DiaKG is a high-quality Chinese dataset for Diabetes knowledge graph.
MAOMaps is a dataset for evaluation of Visual SLAM, RGB-D SLAM and Map Merging algorithms. It contains 40 samples with RGB and depth images, and ground truth trajectories and maps. These 40 samples are joined into 20 pairs of overlapping maps for map merging methods evaluation. The samples were collected using Matterport3D dataset and Habitat simulator.
The Cleft dataset is a collection of ultrasound tongue imaging and audio data, gathered from children with cleft lip and palate by a research speech and language therapist working in a hospital environment.
D-OCC is a large-scale dataset of 5,617 dialogues to enable fine-grained evaluation and analysis of various dialogue systems. It is used to study common grounding in dynamic environments.
The following are all the runs used to generate figures in the paper. Every experiment solves the corresponding high- and low-fidelity model to generate the training, validation, and prediction data.
LIGHT-Quests is an extension of LIGHT, a large-scale crowd-sourced fantasy text-game, to generate a dataset of quests. These contain natural language motivations paired with in-game goals and human demonstrations; completing a quest might require dialogue or actions (or both).
The MT40K dataset for predicting malware threat intelligence is a collection of 40,000 triples generated from 27,354 unique entities and 34 relations. The corpus consists of approximately 1,100 de-identified plain text threat reports written between 2006-2021 and all CVE vulnerability descriptions created between 1990 to 2021. The annotated keyphrases were classified into entities derived from semantic categories defined in malware threat ontologies.
Mouse Brain MRI atlas (both in-vivo and ex-vivo) (repository relocated from the original webpage)
The D3DFACS dataset is a dynamic 3D facial expression data set based on the Facial Action Coding System. It contains Action Unit (AU) sequences from 10 people, with 519 sequences in total. The peak image of each expression sequence has been manually FACS coded by a certified expert.
Clickable heat-map visualizations of the experiments run to quantify the Classic ECN AQM problem and to evaluate the success of the Classic AQM Detection and Fall-back algorithm.
The ClueWeb09 dataset was created to support research on information retrieval and related human language technologies. It consists of about 1 billion web pages in ten languages that were collected in January and February 2009. The dataset is used by several tracks of the TREC conference.
The dataset used for TREC 2017 Dynamic Domain Track consists of two domains: Ebola and New York Times.
Dataset of 5,591 labeled issue tickets. Originally created by Herzig et al. in : "It’s Not a Bug, It’s a Feature: How Misclassification Impacts Bug Prediction" (paper)
Webly-Reference SR dataset is a test dataset for evaluating Ref-SR methods. It has the following advantages:
CPNet dataset has a collection of 25 categories, 2,334 models based on ShapeNetCore, which includes 1,000+ correspondence sets with 104,861 points.