TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Kor-Learner (Korean Learner Corpus)

Kor-Learner is a Korean grammatical error correction (GEC) dataset made from the NIKL learner corpus containing essays written by Korean learners and their grammatical error correction annotations by their tutors in an morpheme-level XML file format. It contains more than 28K sentence pairs.

1 papers0 benchmarksTexts

Kor-Native (Native Korean Corpus)

Kor-Learner is a Korean grammatical error correction (GEC) dataset collected grammatically from two sources, and the correct sentences were read using Google Text-to-Speech(TTS) system. The general public was tasked with dictating grammatically correct sentences and transcribe them. It contains more than 17K sentence pairs.

1 papers0 benchmarksTexts

Kor-Lang8 (Lang-8 Korean Corpus)

Kor-Lang8 is a Korean grammatical error correction (GEC) dataset extracted from the NAIST Lang-8 Learner Corpora by the language label. It contains more than 109K sentence pairs.

1 papers0 benchmarksTexts

pmuBAGE

pmuBAGE (the Benchmarking Assortment of Generated PMU Events) is a dataset that consists of almost 1000 instances of labeled event data to encourage benchmark evaluations on phasor measurement unit (PMU) data analytics. PMU data are challenging to obtain, especially those covering event periods. Nevertheless, power system problems have recently seen phenomenal advancements via data-driven machine learning solutions. A highly accessible standard benchmarking dataset would enable a drastic acceleration of the development of successful machine learning techniques in this field.

1 papers0 benchmarksGraphs

Vident-lab

Vident-lab is a dataset of dental videos with multi-task labels to facilitate further research in relevant video processing applications. The dataset constitutes a low-quality frame, its high-quality counterpart, a teeth segmentation mask, and an inter-frame homography matrix. The homography warps the current frame to the previous frame with respect to the teeth. The dataset has the training, validation, and test sets of 300, 29, and 80 videos, respectively.

1 papers0 benchmarksVideos

ExPUNations

ExPUNations is a humor dataset with such extensive and fine-grained annotations specifically for puns. This dataset is designed for two new tasks namely, explanation generation to aid with pun classification and keyword-conditioned pun generation

1 papers0 benchmarksTexts

Haydn Annotation Dataset

The Haydn Annotation Dataset consists of note onset annotations from 24 experiment participants with varying musical experience. The annotation experiments use recordings from the ARME Virtuoso Strings Dataset.

1 papers0 benchmarksAudio, Music

UJ-CS/Math/Phy

Definitions of jargon/terms in computer science, mathematics, and physics

1 papers0 benchmarksTexts

modified_shemo

A modification on the ShEMO dataset with help of an Automatic Speech Recognition (ASR) system.

1 papers0 benchmarksSpeech, Texts

PKG sample

A random sample from Pubmed Knowledge Graph.

1 papers0 benchmarks

Brazilian Protest

Brazilian Protest is a dataset for event filtering that focuses on protests in multi-modal social media data, with most of the text in Portuguese. The dataset contains 4.5 million tweets, of which 155 thousand are associated with an URL to an uncurated article and 370 thousand have an associated media content (including the media of the uncurated articles).

1 papers0 benchmarksTexts

KITTI-6DoF (KITTI-Six Degrees Of Freedom)

KITTI-6DoF is a dataset that contains annotations for the 6DoF estimation task for 5 object categories on 7,481 frames.

1 papers0 benchmarks3d meshes, 6D, Images

MOET

MOET a dataset consists of gaze data from participants tracking specific objects, annotated with labels and bounding boxes, in crowded real-world videos, for training and evaluating attention decoding algorithms.

1 papers0 benchmarksVideos

NCTE Transcripts

NCTE Transcripts consists of 1,660 45-60 minute long 4th and 5th grade elementary mathematics observations collected by the National Center for Teacher Effectiveness (NCTE) between 2010-2013. The anonymized transcripts represent data from 317 teachers across 4 school districts that serve largely historically marginalized students. The transcripts come with rich metadata, including turn-level annotations for dialogic discourse moves, classroom observation scores, demographic information, survey responses and student test scores.

1 papers0 benchmarksTexts

Spiced

Spiced is a paraphrase dataset of scientific findings annotated for degree of information change. Spiced contains 6,000 scientific finding pairs extracted from news stories, social media discussions, and full texts of original papers.

1 papers0 benchmarksTexts

IMaSC (ICFOSS Malayalam Speech Corpus)

IMaSC is a Malayalam text and speech corpus made available by ICFOSS for the purpose of developing speech technology for Malayalam, particularly text-to-speech. The corpus contains 34,473 text-audio pairs of Malayalam sentences spoken by 8 speakers, totalling in approximately 50 hours of audio.

1 papers0 benchmarksAudio, Texts

PPMR Dataset (pediatric polymicrogyria MRI dataset)

This dataset contains 23 patients in total. Within each patients MRI, not all images show evidence of PMG and some slices are normal. Although the ratio between controls and patients is 3:1, the ratio between normal slices and anomaly slices is around 5:1. Each patient’s brain includes around 150 scans on average.

1 papers0 benchmarks

ODDS (Outlier Detection DataSets (ODDS))

Outliers or anomalies are instances that do not conform to the norm of a dataset. Outlier detection is an important data mining problem that has been researched within diverse research areas and applications domains such as intrusion detection, fraud detection, unusual event detection, disease condition detection etc.

1 papers2 benchmarksTabular, Time series

ProNCI

ProNCI consists of 22.5K proper noun compounds along with their free-form semantic interpretations. ProNCI is 60 times larger than prior noun compound datasets and also includes non-compositional examples.

1 papers0 benchmarksTexts

Greek Parliament Proceedings

Greek Parliament Proceedings is a curated dataset of the Greek Parliament Proceedings that extends chronologically from 1989 up to 2020. It consists of more than 1 million speeches with extensive metadata, extracted from 5,355 parliamentary record files.

1 papers0 benchmarksSpeech
PreviousPage 446 of 1000Next