TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

AraMeter

A dataset to identify the meters of Arabic poems.

0 papers0 benchmarks

Mo2Cap2

A large ground truth training corpus of top-down fisheye images.

0 papers0 benchmarks

Mouse Embryo Tracking Database

The Mouse Embryo Tracking Database is a dataset for tracking mouse embryos. The dataset contains, for each of the 100 examples: (1) the uncompressed frames, up to the 10th frame after the appearance of the 8th cell; (2) a text file with the trajectories of all the cells, from appearance to division (for cells of generations 1 to 3), where a trajectory is a sequence of pairs (center, radius); (3) a movie file showing the trajectories of the cells.

0 papers0 benchmarksImages, Tracking, Videos

Nagoya University Extremely Low-resolution FIR Image Action Dataset

A pedestrian dataset for Person Re-identification.

0 papers0 benchmarksVideos

NERGRIT Corpus

NERGRIT involves machine learning based NLP Tools and a corpus used for Indonesian Named Entity Recognition, Statement Extraction, and Sentiment Analysis.

0 papers0 benchmarksTexts

NetiLook

A large-scale clothing dataset named NetiLook to discover netizen-style comments.

0 papers0 benchmarks

Neural Code Search Evaluation Dataset

Neural-Code-Search-Evaluation-Dataset presents an evaluation dataset consisting of natural language query and code snippet pairs, with the hope that future work in this area can use this dataset as a common benchmark.

0 papers0 benchmarks

PAF Benchmark

Introduce three new neuromorphic vision datasets recorded by a novel neuromorphic vision sensor named Dynamic Vision Sensors (DVS).

0 papers0 benchmarks

NSMC (Naver Sentiment Movie Corpus)

This is a movie review dataset in the Korean language. Reviews were scraped from Naver Movies.

0 papers0 benchmarksTexts

NYC3DCars

A vehicle detection database for vision tasks set in the real world.

0 papers0 benchmarksImages

OffComBR (Offensive Comments in the Brazilian Web)

Offensive comments obtained from Brazilian website.

0 papers0 benchmarks

One Million Posts Corpus

An annotated data set consisting of user comments posted to an Austrian newspaper website (in German language).

0 papers0 benchmarks

OpenSurfaces

OpenSurfaces is a large database of annotated surfaces created from real-world consumer photographs. The framework used for the annotation process draws on crowdsourcing to segment surfaces from photos, and then annotate them with rich surface properties, including material, texture and contextual information.

0 papers0 benchmarks3D, Images

Opinosis

This dataset contains sentences extracted from user reviews on a given topic. Example topics are “performance of Toyota Camry” and “sound quality of ipod nano”, etc. In total there are 51 such topics with each topic having approximately 100 sentences (on average). The reviews were obtained from various sources – Tripadvisor (hotels), Edmunds.com (cars) and Amazon.com (various electronics). This dataset was used for the following automatic text summarization project .

0 papers0 benchmarks

Spherical-Navi

A novel 360◦ fisheye panoramas dataset, i.e., the Spherical-Navi image dataset is collected, with a unique labeling strategy enabling automatic generation of an arbitrary number of negative samples (wrong heading direction).

0 papers0 benchmarksImages

Plaintext Jokes

There are about 208 000 jokes in this database scraped from three sources.

0 papers0 benchmarksTexts

prachathai-67k

The prachathai-67k dataset was scraped from the news site Prachathai excluding articles with less than 500 characters of body text (mostly images and cartoons). It contains 67,889 articles with 51,797 tags from August 24, 2004 to November 15, 2018.

0 papers0 benchmarksTexts

QuAIL (Question Answering for Artificial Intelligence)

A new kind of question-answering dataset that combines commonsense, text-based, and unanswerable questions, balanced for different genres and reasoning types. Reasoning type annotation for 9 types of reasoning: temporal, causality, factoid, coreference, character properties, their belief states, subsequent entity states, event durations, and unanswerable. Genres: CC license fiction, Voice of America news, blogs, user stories from Quora 800 texts, 18 questions for each (~14K questions).

0 papers0 benchmarksTexts

RAF-ML (Real-world Affective Faces Multi Label)

Real-world Affective Faces Multi Label (RAF-ML) is a multi-label facial expression dataset with around 5K great-diverse facial images downloaded from the Internet with blended emotions and variability in subjects' identity, head poses, lighting conditions and occlusions. During annotation, 315 well-trained annotators are employed to ensure each image can be annotated enough independent times. And images with multi-peak label distribution are selected out to constitute the RAF-ML.

0 papers0 benchmarksImages

RGB-D Object dataset

The dataset contains 300 objects organized into 51 categories and has been made publicly available to the research community so as to enable rapid progress based on this promising technology.

0 papers0 benchmarks
PreviousPage 637 of 1000Next