TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Dataset for Post-OCR text correction in Sanskrit

This dataset contains around 218K sentences, with 1.5 million words, from 30 different books designed for Post-OCR text correction.

1 papers0 benchmarksImages, Texts

INDRA (INdian Dataset for RoAd crossing)

INDRA is a dataset capturing videos of Indian roads from the pedestrian point-of-view. INDRA contains 104 videos comprising of 26k 1080p frames, each annotated with a binary road crossing safety label and vehicle bounding boxes.

1 papers0 benchmarksVideos

MM-Locate-News

MM-Locate-News is a dataset for location estimation of news. It consists of 6395 news articles covering 237 cities and 152 countries across all continents as well as multiple domains such as health, environment, and politics. The dataset is collected in a weakly-supervised manner, and multiple data cleaning steps are applied to remove articles with potential inaccurate geolocation information. The acquired dataset addresses drawbacks of other datasets such as BreakingNews as it considers multimodal content of news to label the corresponding location.

1 papers0 benchmarksImages, Texts

Nlakh

Nlakh is a dataset for Musical Instrument Retrieval. It is a combination of the NSynth dataset, which provides a large number of instruments, and the Lakh dataset, which provides multi-track MIDI data.

1 papers0 benchmarksAudio, Music

YM2413-MDB

YM2413-MDB is an 80s FM video game music dataset with multi-label emotion annotations. It includes 669 audio and MIDI files of music from Sega and MSX PC games in the 80s using YM2413, a programmable sound generator based on FM. The collected game music is arranged with a subset of 15 monophonic instruments and one drum instrument.

1 papers0 benchmarksMusic

PaintNet

PaintNet is a dataset for learning robotic spray painting of free-form 3D objects. PaintNet includes more than 800 object meshes and the associated painting strokes collected in a real industrial setting.

1 papers0 benchmarks3D, 3d meshes, 6D

ec-darkpattern

ec-darkpattern is a dataset for dark pattern detection and prepared its baseline detection performance with state-of-the-art machine learning methods. The original dataset was obtained from Mathur et al.’s study in 2019 [11kScale], which consists of 1,818 dark pattern texts from shopping sites. Negative samples, i.e., non-dark pattern texts, by retrieving texts from the same websites as Mathur et al.'s dataset.

1 papers0 benchmarksTexts

kaggle stroke Prediction competition

It is a competition on kaggle with stroke Prediction, which is heavily imbalanced.

1 papers0 benchmarksMedical, Tables

Virtuoso Strings

Virtuoso Strings is a dataset for soft onsets detection for string instruments. It consists of over 144 recordings of professional performances of an excerpt from Haydn's string quartet Op. 74 No. 1 Finale, each with corresponding individual instrumental onset annotations.

1 papers0 benchmarksAudio, Music

CIPM evaluation data (Evaluation data for Continuous Integration of Performance Models (CIPM))

The main goal of the Continuous Integration of Performance Models (CIPM) is to enable an accurate architecture-based performance prediction at each point of the systems development life cycle. For this goal, the CIPM approach continuously updates the architecture-level performance models of a software system according to observed changes at development time and at operation time.

1 papers0 benchmarks

Comet

Comet is a dataset which contains 11.5k user-assistant dialogs (totalling 103k utterances), grounded in simulated personal memory graphs.

1 papers0 benchmarksTexts

GF-PA66 3D XCT (latest) (Glass fiber-reinforced polyamide 66 3D X-ray Computed Tomography)

Stack of 2D gray images of glass fiber-reinforced polyamide 66 (GF-PA66) 3D X-ray Computed Tomography (XCT) specimen.

1 papers0 benchmarks3D, Images

IEEE CIS Fraud Detection (IEEE CIS Fraud Detection (Vesta))

The Vesta dataset was released for use in the IEEE CIS Fraud Detection competition. It contains 590,540 card transactions, 20,663 of which are fraudulent (3.5%). Each transaction has 431 features (400 numerical, 31 categorical), along with the relative timestamp and a label of whether it was fraudulent or legitimate. For anonymization purposes, the names of the identity features have been masked, along with the names of the extra features engineered by Vesta.

1 papers0 benchmarks

Padding Ain't Enough: Assessing the Privacy Guarantees of Encrypted DNS – Web Scans

This dataset contains the main data set of our FOCI 2020 paper "Padding Ain’t Enough: Assessing the Privacy Guarantees of Encrypted DNS".

1 papers0 benchmarks

Padding Ain't Enough: Assessing the Privacy Guarantees of Encrypted DNS – Subpage-Agnostic Domain Classification Firefox

This dataset contains one part for the "Subpage-Agnostic Domain Classification" section of our FOCI 2020 paper "Padding Ain’t Enough: Assessing the Privacy Guarantees of Encrypted DNS".

1 papers0 benchmarks

Padding Ain't Enough: Assessing the Privacy Guarantees of Encrypted DNS – Subpage-Agnostic Domain Classification Tor Browser

This dataset contains the second part of the "Subpage-Agnostic Domain Classification" section of our FOCI 2020 paper "Padding Ain’t Enough: Assessing the Privacy Guarantees of Encrypted DNS".

1 papers0 benchmarks

Datasets for automatic acoustic identification of insects (Orthoptera and Cicadidae)

This dataset contains recordings of 32 sound producing insect species with a total 335 files and a length of 57 minutes. The dataset was compiled for training neural networks to automatically identify insect species while comparing adaptive, waveform-based frontends to conventional mel-spectrogram frontends for audio feature extraction. This work will be submitted for publication in the future and this dataset can be used to replicate the results, as well as other uses. The scripts for audio processing and the machine learning implementations will be published on Github.

1 papers0 benchmarksAudio, Biology, Environment

SummZoo

SummZoo, a benchmark consists of 8 diverse summarization tasks with multiple sets of few-shot samples for each task, covering both monologue and dialogue domains.

1 papers0 benchmarksTexts

IDK-MRC

IDK-MRC is an Indonesian Machine Reading Comprehension (MRC) dataset consists of more than 10K questions in total with over 5K unanswerable questions with diverse question types.

1 papers0 benchmarksTexts

CCSE (Chinese Character Stroke Extraction)

Chinese Character Stroke Extraction (CCSE) is a benchmark containing two large-scale datasets: Kaiti CCSE (CCSE-Kai) and Handwritten CCSE (CCSE-HW). It is designed for stroke extraction problems.

1 papers0 benchmarksImages, Texts
PreviousPage 445 of 1000Next