TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Visuomotor affordance learning (VAL) robot interaction dataset

This data contains about 2500 trajectories (with images and actions) of a Sawyer robot interacting with various objects.

1 papers0 benchmarksActions, Images, Videos

MLQuestions

MLQuestions is a domain-adaptation dataset for the machine learning domain containing 50K unaligned passages and 35K unaligned questions, and 3K aligned passage and question pairs.

1 papers0 benchmarksTexts

ARC Ukiyo-e Faces

ARC Ukiyo-e Faces is a large-scale (>10k paintings, >20k faces) Ukiyo-e dataset with coherent semantic labels and geometric annotations through augmenting and organizing existing datasets with automatic detection.

1 papers0 benchmarksImages

Quo Vadis, Open Source? (Quo Vadis, Open Source? The Limits of Open Source Growth)

This is an complete set of the data we collected and analyzed in our study "Quo Vadis, Open Source? The Limits of Open Source Growth". Please see our GitHub repository for details and tool chain.

1 papers0 benchmarksTime series

Webis-ConcluGen-21

Webis-ConcluGen-21 is a large-scale corpus of 136,996 samples of argumentative texts and their conclusions used for the task of generating informative conclusions.

1 papers0 benchmarksTexts

CTFW

CTFW is a large annotated procedural text dataset in the cybersecurity domain (3154 documents). It is used to generate flow graphs from procedural texts.

1 papers0 benchmarksGraphs, Texts

Onmiglot

1 papers0 benchmarks

LSEC (Live Stream E-Commerce)

The LSEC (Live Stream E-Commerce) dataset has two subsets: LSEC-Small and LSEC-Large. It is a dataset for studying E-commerce transactions in the context of live streams, where the streames are talking about products while interacting with their audience. The dataset consists of interaction information among streamers, users, and products.

1 papers0 benchmarksGraphs

Unsplash2K

Unsplash2K is high-resolution image dataset with 2K resolution. Unsplash2K dataset is crawled from unsplash. Unsplash2K dataset contains 498 high-resolution images and corresponding low-resolution images which are downsampled by bicubic downsamling for x2, x4, x8 scale. Unsplash2K contains diverse contents such as animals, architectures and flowers.

1 papers0 benchmarksImages

Emol news articles and comments

The dataset provides News articles obtained from emol.cl including their content, title and all the comments it received in JSON format

1 papers0 benchmarks

FastZIP Data (FastZIP Dataset and Code)

Structure of code/data folders and how to use them fastzip-code

1 papers0 benchmarksEnvironment

TESTIMAGES

A collection of photographic and synthetic images intended for analysis of image processing techniques and quality assessment of displays.

1 papers0 benchmarksImages

Notre-Dame Cathedral Fire

Number of images: 1,657 images during or after the fire

1 papers0 benchmarksImages

S_B_D (Synthetic Barcode Dataset)

100,000 LR synthetic barcode datasets along with their corresponding bounding boxes ground truth masks.

1 papers0 benchmarks

Date Estimation in the Wild

~1M Flickr images from the XX century-aged from the 1910s to 1990s. Dataset was introduced by Müller et al. and can be found https://www.radar-service.eu/radar/en/dataset/tJzxrsYUkvPklBOw

1 papers0 benchmarksImages, Ranking

Symmetric Solids

This is a pose estimation dataset, consisting of symmetric 3D shapes where multiple orientations are visually indistinguishable. The challenge is to predict all equivalent orientations when only one orientation is paired with each image during training (as is the scenario for most pose estimation datasets). In contrast to most pose estimation datasets, the full set of equivalent orientations is available for evaluation.

1 papers0 benchmarksImages

Evidence-based Factual Error Correction

Intermediate annotations from the FEVER dataset that describe original facts extracted from Wikipedia and the mutations that were applied, yielding the claims in FEVER.

1 papers0 benchmarksTexts

Bus Trajectory Dataset

This dataset contains the bus trajectory dataset collected by 6 volunteers who were asked to travel across the sub-urban city of Durgapur, India, on intra-city buses (route name: 54 Feet). During the travel, the volunteers captured sensor logs through an Android application installed on COTS smartphones.

1 papers0 benchmarksActions, Environment, Stereo

MARS-DL

MARS dataset processed with our re-Detect and Link (DL) module.

1 papers0 benchmarksImages

DukeMTMC-VideoReID-DL

DukeMTMC-VideoReID-DL processed with our re-Detect and Link (DL) module.

1 papers0 benchmarks
PreviousPage 396 of 1000Next