TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

SCIAN (SCIAN Gold-standard for Morphological Sperm Analysis)

Dataset of sperm head images with expert-classification labels. The dataset contains 1854 sperm head images obtained from six semen smears and classified by three Chilean referent domain experts according to World Health Organization (WHO) criteria, in one of the following classes: normal, tapered, pyriform, small and amorphous. This gold-standard is aimed for use in evaluating and comparing not only known techniques, but also future improvements to present approaches for classification of human sperm heads for semen analysis.

1 papers0 benchmarksImages, Medical

A collection of LFR benchmark graphs

This dataset is a collection of undirected and unweighted LFR benchmark graphs as proposed by Lancichinetti et al. [1]. We generated the graphs using the code provided by Santo Fortunato on his personal website [2], embedded in our evaluation framework [3], with two different parameter sets. Let N denote the number of vertices in the network, then

1 papers0 benchmarks

Volunteer task execution events in Galaxy Zoo and The Milky Way citizen science projects

Context of the data sets The Zooniverse platform (www.zooniverse.org) has successfully built a large community of volunteers contributing to citizen science projects. Galaxy Zoo and the Milky Way Project were hosted there.

1 papers0 benchmarksActions, Tabular, Time series

Motor Imagery dataset

From dataset repository for "2020 International BCI Competition": https://osf.io/pq7vb/?view_only=08e7108d89fd42bab2adbd6b98fb683d

1 papers0 benchmarks

Error Grids for multi-fidelity benchmark functions in mf2

Provide:

1 papers0 benchmarksTabular

KITTI'15 MSplus (KITTI'15 Instance Motion Segmentation - extension)

Extension of the official KITTI'15 dataset. independently moving instance segmentation ground truth to cover all moving objects, not just a selection of cars and vans.

1 papers0 benchmarksImages

NON-LINEAR PHASE NOISE MITIGATION OVER SYSTEMS USING CONSTELLATION SHAPING: EXPERIMENTAL DATASET (Dario Pilori)

This dataset contains the full set of experimental waveforms that were used to produce the article "Non-Linear Phase Noise Mitigation over Systems using Constellation Shaping", published in the Journal of Lightwave Technology with DOI: 10.1109/JLT.2019.2917308.

1 papers0 benchmarks

Typography-MNIST

Typography-MNIST is a dataset comprising of 565,292 MNIST-style grayscale images representing 1,812 unique glyphs in varied styles of 1,355 Google-fonts. The glyph-list contains common characters from over 150 of the modern and historical language scripts with symbol sets, and each font-style represents varying subsets of the total unique glyphs. The dataset has been developed as part of the Cognitive Type project which aims to develop eye-tracking tools for real-time mapping of type to cognition and to create computational tools that allow for the easy design of typefaces with cognitive properties such as readability.

1 papers0 benchmarks

VaccineLies

A Natural Language Resource for Learning to Recognize Misinformation about the COVID-19 and HPV Vaccines.

1 papers0 benchmarksTexts

EUCA dataset

EUCA dataset description Associated Paper: EUCA: the End-User-Centered Explainable AI Framework

1 papers0 benchmarksTabular

Malnutrition data (malnutrition data from UN)

The malnutrition data, from the United Nations Children's Fund data warehouse, include two variables, stunted growth and the prevalence of low birth weight, collected in 77 countries from 1985 to 2019. Stunted growth is defined as the proportion of newborns aging from 0 to 59 months with a low height-for-age measurement (below two standard deviations). The stunted growth data represent a point sparseness case with 4-23 recordings per nation. The low birth weight data are a partial sparseness case, with recordings during 2000-2015 only.

1 papers0 benchmarks

MuMiN-small

This is the small version of the MuMiN dataset.

1 papers2 benchmarksGraphs, Images, Texts

MuMiN-medium

This is the medium version of the MuMiN dataset.

1 papers2 benchmarksGraphs, Images, Texts

MuMiN-large

This is the large version of the MuMiN dataset.

1 papers2 benchmarksGraphs, Images, Texts

iFLYTEK

iFLYTEK and ChangGuang Satellite jointly held the challenge of extracting cultivated land from high-resolution remote sensing images.

1 papers0 benchmarks

MuVi (MusicVideos)

A dataset of music videos with continuous valence/arousal ratings as well as emotion tags.

1 papers0 benchmarksMusic, Videos

GF-PA66 3D XCT (Glass fiber-reinforced polyamide 66 (GF-PA66) 3D X-ray Computed Tomography (XCT)))

Stack of 2D gray images of glass fiber-reinforced polyamide 66 (GF-PA66) 3D X-ray Computed Tomography (XCT) specimen.

1 papers1 benchmarks3D, Images

Icon645

Icon645 is a large-scale dataset of icon images that cover a wide range of objects:

1 papers0 benchmarksImages

AirSim Stereo Synthetic Dataset

Synthetic Dataset created in AirSim

1 papers0 benchmarks

TraVLR

TraVLR is a synthetic dataset comprising four visio-linguistic reasoning tasks. Each example encodes the scene bimodally such that either modality can be dropped during training/testing with no loss of relevant information. TraVLR's training and testing distributions are also constrained along task-relevant dimensions, enabling the evaluation of out-of-distribution generalisation.

1 papers0 benchmarksTexts
PreviousPage 420 of 1000Next