TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

SPI dataset

The SPI dataset consists of force-controlled industrial robot data for training shadow program inversion (SPI) models.

1 papers0 benchmarks

NAVER LABS Localization Datasets

The NAVER LABS localization datasets are 5 new indoor datasets for visual localization in challenging real-world environments. They were captured in a large shopping mall and a large metro station in Seoul, South Korea, using a dedicated mapping platform consisting of 10 cameras and 2 laser scanners. In order to obtain accurate ground truth camera poses, we used a robust LiDAR SLAM which provides initial poses that are then refined using a novel structure-from-motion based optimization. The datasets are provided in the kapture format and contain about 130k images as well as 6DoF camera poses for training and validation. We also provide sparse Lidar-based depth maps for the training images. The poses of the test set are withheld to not bias the benchmark.

1 papers0 benchmarks

behavioral observation data entry apps

In this repository, we provide the set-up files and output files of 5 behavioral observation data entry applications. These applications allow observers to collect animal behavior data on a handheld computer (phone/tablet).

1 papers0 benchmarks

BigCQ

BigCQ is a dataset of Competency Question templates paired with SPARQL-OWL query templates. These represent templates of ontology requirements formalizations which are then translated into SPARQL-OWL query language used to query T-Box level of ontologies. Thus, such a dataset can be used in various scenarios regarding ontology authoring:

1 papers0 benchmarks

NewsMTSC

NewsMTSC is a dataset for target-dependent sentiment classification (TSC) on news articles reporting on policy issues. The dataset consists of more than 11k labeled sentences, which we sampled from news articles from online US news outlets.

1 papers0 benchmarksTexts

ZuBuD (Zurich Buildings Database)

The goal of the ZuBuD Image Database is to share image data sets with researcheres around the world. To facilitate this, we have created this site, which contains over 1005 images about Zurich city building. The detail information about the database can be found on our Technical Report:TR-​260.

1 papers0 benchmarks

The RBO Dataset of Articulated Objects and Interactions

The RBO dataset of articulated objects and interactions is a collection of 358 RGB-D video sequences (67:18 minutes) of humans manipulating 14 articulated objects under varying conditions (light, perspective, background, interaction). All sequences are annotated with ground truth of the poses of the rigid parts and the kinematic state of the articulated object (joint states) obtained with a motion capture system. We also provide complete kinematic models of these objects (kinematic structure and three-dimensional textured shape models). In 78 sequences the contact wrenches during the manipulation are also provided.

1 papers0 benchmarks3d meshes, Point cloud, RGB-D, Time series, Videos

Clarkson Fingerprint Generator

Clarkson Fingerprint Generator consists of a dataset of 50K synthetically generated fingerprints.

1 papers0 benchmarksImages

scb_name_length_data_Sweden_Stockholm_2019 (SCB's Name Length Data in Sweden and Stockholm)

Appendix A in this paper contains a real-world name length data for the whole of Sweden as well as Stockholm Municipality (Swedish: Stockholms kommun) as of 31 December 2019. It excludes names that either belong to people with protected identities or are suspiciously incorrect due to errors in petition. But these excluded numbers are low and should not matter for statistical purposes.

1 papers0 benchmarks

TabStructDB

In ICDAR-17, a Page-Object Detection (POD) competition was organized where the task was to identify page objects in documents which includes tables, figures and equations in document. The dataset was composed of 2417 images in total, where 1600 images were used for training, while the rest of the 817 images were used for testing. We are introducing a new table structure recognition dataset, TabStructDB, where we labeled each tabular region present in the ICDAR-17 POD dataset with table structure information comprising of the row and column information.

1 papers0 benchmarksImages

WikiBioCTE

WikiBioCTE is a dataset for controllable text edition based on the existing dataset WikiBio (originally created for table-to-text generation). In the task of controllable text edition the input is a long text, a question, and a target answer, and the output is a minimally modified text, so that it fits the target answer. This task is very important in many situations, such as changing some conditions, consequences, or properties in a legal document, or changing some key information of an event in a news text.

1 papers0 benchmarksTexts

Dataset for: "It is just a flu: Assessing the Effect of Watch History on YouTube's Pseudoscientific Video Recommendations"

The dataset consists of three files: the metadata, comments, and captions of the ground-truth dataset videos collected and manually reviewed in this paper.

1 papers0 benchmarks

MacaquePose

MacaquePose is an animal pose estimation dataset containing pictures of macaque monkeys and manually labeled annotations on them.

1 papers2 benchmarksImages

Vinegar Fly

Vinegar Fly is a pose estimation dataset for fruit flies.

1 papers2 benchmarksImages

Desert Locust

Desert Locus is a animal pose estimation dataset for desert locuses.

1 papers2 benchmarks

USM-SED

USM-SED is a dataset for polyphonic sound event detection in urban sound monitoring use-cases. Based on isolated sounds taken from the FSD50k dataset, 20,000 polyphonic soundscapes are synthesized with sounds being randomly positioned in the stereo panorama using different loudness levels.

1 papers0 benchmarksAudio

CEREC (Corpus for Entity Resolution in Email Conversations)

CEREC is a large scale corpus for entity resolution in email conversations. The corpus consists of 6001 email threads from the Enron Email Corpus containing 36,448 email messages and 60,383 entity coreference chains. The annotation is carried out as a two-step process with minimal manual effort.

1 papers0 benchmarksTexts

Custom FINNgers

A dataset with 3200 images (200 for each number quantity on each hand).

1 papers5 benchmarks

Sentinel 2 manually extracted deep water spectra with high noise levels and sunglint

This dataset includes 2.133.324 reflectance water spectra which were manually extracted by visual observation from 30 Sentinel 2 level 1C satellite images. The spectra were extracted from deep water areas with high noise levels and sunglint. The Sentinel 2 images depicted 2 tiles of the same orbit and were collected in 2016 (2 images), 2017 (19 images) and 2018 (9 images). The images contain 13 bands, 3 with 60 m spatial resolution, 4 with 10 m spatial resolution and 6 with 20 m spatial resolution. Before the spectra extraction, the bands with spatial resolution 10 and 20 m were resampled to 60 m and then the images were cropped in order to remove the land and depict optically homogenous sea regions. A figure depicting the location of the Sentinel 2 tiles (white polygons (1,2)) and the cropped tiles (red polygons (3,4)) is included in this folder. A figure depicting example scenes from which spectra were obtained through regions of interest (rois) is included as well. The spectra are s

1 papers0 benchmarks

TexRel

Green family of datasets for emergent communications on relations.

1 papers0 benchmarksImages, Texts
PreviousPage 394 of 1000Next