TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

ReVerb45K

Open KB canonicalization dataset. ReVerb45K increases the entity number to 7.5K and has 45K triples in total. ReVerb45K extract a source sentence for each triple from ClueWeb09

1 papers0 benchmarks

APPBENCH (SVBRDF Database Bonn)

A database of 56 high quality fabric material measurements, provided as carefully calibrated rectified HDR images, together with SVBRDF fits. Used in the Fabric Appearance Challange.

1 papers0 benchmarksImages

RUSHOLD (Roman Urdu Hate Speech and Offensive Language Dataset)

RUHSOLD is hate speech and offensive language dataset in Roman Urdu. The dataset contains over 10 thousand tweets that are hand labelled into the following categories: 1) Abusive/Offensive 2) Untargeted 3) Sexism 4) Religious 5) Neutral

1 papers0 benchmarks

B-T4SA

1 papers1 benchmarksImages, Texts

Kinect-WSJ

Kinect-WSJ is a multichannel, multispeaker, reverberated, noisy dataset which extends the WSJ0-2mix singlechannel, non-reverberated, noiseless dataset to the strong reverberation and noise conditions and the Kinect-like microphone array geometry used in CHiME-5.

1 papers0 benchmarksAudio, Speech

AbstRCT - Neoplasm

The AbstRCT dataset consists of randomized controlled trials retrieved from the MEDLINE database via PubMed search. The trials are annotated with argument components and argumentative relations.

1 papers3 benchmarksTexts

OTEANNv3

This dataset contains orthographic samples of words in 19 languages (ar, br, de, en, eno, ent, eo, es, fi, fr, fro, it, ko, nl, pt, ru, sh, tr, zh). Each sample contains two text features: a Word (the textual representation of the word according to its orthography) and a Pronunciation (the highest-surface IPA pronunciation of the word as pronunced in its language).

1 papers0 benchmarks

Maintenance of Wakefulness Test (MWT) recordings

Maintenance of Wakefulness Test (MWT) is a dataset of recordings with microsleep episodes and drowsiness.

1 papers0 benchmarksEEG

darpa_sd2_perovskites

Included in this content:

1 papers0 benchmarks

Deep Thermal Imaging Dataset

The Deep Thermal Imaging dataset consists of two main datasets:

1 papers0 benchmarksImages

Fongbe audio (Fongbe dataset)

Fongbe Data collected by Fréjus A. A LALEYE

1 papers1 benchmarksAudio

Lens Flare Dataset

The Lens Flare dataset is an internal dataset for Flare Spot detection used in the paper "Automatic Flare Spot Artifact Detection and Removal in Photographs" by Patricia Vitoria and Coloma Ballester.

1 papers0 benchmarks

SARA motion (Synthetic Actors and Real Actions)

Sara motion is a 3D motion dataset, named Synthetic Actors and Real Actions (SARA), for training a model to produce motion embeddings suitable for reasoning about motion similarity.

1 papers0 benchmarks3D, Videos

NTU RGB+D 120 motion similarity

Motion similarity annotations for NTU RGB+D 120 dataset to evaluate motion similarity in the real world.

1 papers0 benchmarksImages, Videos

BU-BIL (Boston University Biomedical Image Library)

BU-BIL is an image library which includes six datasets that represent three imaging modalities and six object types. Providers of the datasets are instructed to choose images that capture the various environmental conditions and imaging noise that arose in their studies. These experts are asked to then select objects from those images that reflect the natural diversity of shape and appearances that these objects can exhibit. The image subregions containing the identified objects are cropped to create the image library. The outcome was a library with 305 objects from 235 images. Authors verify by visual inspection that the image library includes a variety of object appearances, backgrounds, and properties distinguishing objects from the background.

1 papers0 benchmarks

MTA-KDD'19 (Malware Traffic Analysis Knowledge Dataset 2019)

Malware Traffic Analysis Knowledge Dataset 2019 (MTA-KDD'19) is an updated and refined dataset specifically tailored to train and evaluate machine learning based malware traffic analysis algorithms. To generate it, that authors started from the largest databases of network traffic captures available online, deriving a dataset with a set of widely-applicable features and then cleaning and preprocessing it to remove noise, handle missing data and keep its size as small as possible. The resulting dataset is not biased by any specific application (although specifically addressed to machine learning algorithms), and the entire process can run automatically to keep it updated.

1 papers0 benchmarks

Cuff-Less Blood Pressure Estimation (Cuff-Less Blood Pressure Estimation. Pre-processed and cleaned vital signals for cuff-less BP estimation.)

Data Set Information: The main goal of this data set is providing clean and valid signals for designing cuff-less blood pressure estimation algorithms. The raw electrocardiogram (ECG), photoplethysmograph (PPG), and arterial blood pressure (ABP) signals are originally collected from the physionet.org and then some preprocessing and validation performed on them. (For more information about the process please refer to our paper)

1 papers0 benchmarksMedical

POTUS Corpus

The POTUS Corpus is a Database of Weekly Addresses for the Study of Stance in Politics and Virtual Agents.

1 papers0 benchmarksAudio, Videos

ImageNet VIPriors subset

The training and validation data are subsets of the training split of the Imagenet 2012. The test set is taken from the validation split of the Imagenet 2012 dataset. Each data set includes 50 images per class.

1 papers0 benchmarksImages

LIFULL HOME'S

The National Institute of Informatics provides LIFULL HOME'S Dataset to researchers, which was offered by LIFULL Co., Ltd. for promoting research in informatics and the related fields.

1 papers0 benchmarksImages
PreviousPage 384 of 1000Next