19,997 machine learning datasets
19,997 dataset results
The SPI dataset consists of force-controlled industrial robot data for training shadow program inversion (SPI) models.
The NAVER LABS localization datasets are 5 new indoor datasets for visual localization in challenging real-world environments. They were captured in a large shopping mall and a large metro station in Seoul, South Korea, using a dedicated mapping platform consisting of 10 cameras and 2 laser scanners. In order to obtain accurate ground truth camera poses, we used a robust LiDAR SLAM which provides initial poses that are then refined using a novel structure-from-motion based optimization. The datasets are provided in the kapture format and contain about 130k images as well as 6DoF camera poses for training and validation. We also provide sparse Lidar-based depth maps for the training images. The poses of the test set are withheld to not bias the benchmark.
In this repository, we provide the set-up files and output files of 5 behavioral observation data entry applications. These applications allow observers to collect animal behavior data on a handheld computer (phone/tablet).
BigCQ is a dataset of Competency Question templates paired with SPARQL-OWL query templates. These represent templates of ontology requirements formalizations which are then translated into SPARQL-OWL query language used to query T-Box level of ontologies. Thus, such a dataset can be used in various scenarios regarding ontology authoring:
NewsMTSC is a dataset for target-dependent sentiment classification (TSC) on news articles reporting on policy issues. The dataset consists of more than 11k labeled sentences, which we sampled from news articles from online US news outlets.
The goal of the ZuBuD Image Database is to share image data sets with researcheres around the world. To facilitate this, we have created this site, which contains over 1005 images about Zurich city building. The detail information about the database can be found on our Technical Report:TR-260.
The RBO dataset of articulated objects and interactions is a collection of 358 RGB-D video sequences (67:18 minutes) of humans manipulating 14 articulated objects under varying conditions (light, perspective, background, interaction). All sequences are annotated with ground truth of the poses of the rigid parts and the kinematic state of the articulated object (joint states) obtained with a motion capture system. We also provide complete kinematic models of these objects (kinematic structure and three-dimensional textured shape models). In 78 sequences the contact wrenches during the manipulation are also provided.
Clarkson Fingerprint Generator consists of a dataset of 50K synthetically generated fingerprints.
Appendix A in this paper contains a real-world name length data for the whole of Sweden as well as Stockholm Municipality (Swedish: Stockholms kommun) as of 31 December 2019. It excludes names that either belong to people with protected identities or are suspiciously incorrect due to errors in petition. But these excluded numbers are low and should not matter for statistical purposes.
In ICDAR-17, a Page-Object Detection (POD) competition was organized where the task was to identify page objects in documents which includes tables, figures and equations in document. The dataset was composed of 2417 images in total, where 1600 images were used for training, while the rest of the 817 images were used for testing. We are introducing a new table structure recognition dataset, TabStructDB, where we labeled each tabular region present in the ICDAR-17 POD dataset with table structure information comprising of the row and column information.
WikiBioCTE is a dataset for controllable text edition based on the existing dataset WikiBio (originally created for table-to-text generation). In the task of controllable text edition the input is a long text, a question, and a target answer, and the output is a minimally modified text, so that it fits the target answer. This task is very important in many situations, such as changing some conditions, consequences, or properties in a legal document, or changing some key information of an event in a news text.
The dataset consists of three files: the metadata, comments, and captions of the ground-truth dataset videos collected and manually reviewed in this paper.
MacaquePose is an animal pose estimation dataset containing pictures of macaque monkeys and manually labeled annotations on them.
Vinegar Fly is a pose estimation dataset for fruit flies.
Desert Locus is a animal pose estimation dataset for desert locuses.
USM-SED is a dataset for polyphonic sound event detection in urban sound monitoring use-cases. Based on isolated sounds taken from the FSD50k dataset, 20,000 polyphonic soundscapes are synthesized with sounds being randomly positioned in the stereo panorama using different loudness levels.
CEREC is a large scale corpus for entity resolution in email conversations. The corpus consists of 6001 email threads from the Enron Email Corpus containing 36,448 email messages and 60,383 entity coreference chains. The annotation is carried out as a two-step process with minimal manual effort.
A dataset with 3200 images (200 for each number quantity on each hand).
This dataset includes 2.133.324 reflectance water spectra which were manually extracted by visual observation from 30 Sentinel 2 level 1C satellite images. The spectra were extracted from deep water areas with high noise levels and sunglint. The Sentinel 2 images depicted 2 tiles of the same orbit and were collected in 2016 (2 images), 2017 (19 images) and 2018 (9 images). The images contain 13 bands, 3 with 60 m spatial resolution, 4 with 10 m spatial resolution and 6 with 20 m spatial resolution. Before the spectra extraction, the bands with spatial resolution 10 and 20 m were resampled to 60 m and then the images were cropped in order to remove the land and depict optically homogenous sea regions. A figure depicting the location of the Sentinel 2 tiles (white polygons (1,2)) and the cropped tiles (red polygons (3,4)) is included in this folder. A figure depicting example scenes from which spectra were obtained through regions of interest (rois) is included as well. The spectra are s
Green family of datasets for emergent communications on relations.