TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

A Datacube for the analysis of wildfires in Greece

This dataset is meant to be used to develop models for next-day fire hazard forecasting in Greece. It contains data from 2009 to 2020 at a 1km x 1km x 1 daily grid.

1 papers0 benchmarksEnvironment, Videos

CLUES (Constrained Language Understanding Evaluation Standard)

CLUES (Constrained Language Understanding Evaluation Standard) is a benchmark for evaluating the few-shot learning capabilities of NLU models.

1 papers0 benchmarksTexts

WWU DUNEuro reference data set (The WWU DUNEuro reference data set for combined EEG/MEG source analysis)

The provided dataset consists of high-quality realistic head models and combined EEG/MEG data which can be used for state-of-the-art methods in brain research, such as modern finite element methods (FEM) to compute the EEG/MEG forward problems using the software toolbox DUNEuro (http://duneuro.org).

1 papers0 benchmarks3d meshes, EEG, Medical

VSLID (Very Small Lego Image Dataset)

VSLID stands for Very Small Lego Image Dataset. It has a bit over 1800 images of piles of LEGO bricks of 85 different types. There are between 1 and 10 bricks per image. Backgrounds and lighting conditions vary. All images are annotated with a list of the visible bricks. The images can have two resolutions, so rescaling them is recommended before usage.

1 papers0 benchmarks

RIKEN Microstructural Imaging Metadatabase

The RIKEN Microstructural Imaging Metadatabase is a semantic web-based imaging database in which image metadata are described using the Resource Description Framework (RDF) and detailed biological properties observed in the images can be represented as Linked Open Data. The metadata are used to develop a large-scale imaging viewer that provides a straightforward graphical user interface to visualise a large microstructural tiling image at the gigabyte level.

1 papers0 benchmarks

SDSS Galaxies (SDSS galaxies as imaged by DESI)

This is a dataset of 306,006 galaxies whose coordinates are taken from the Sloan Digital Sky Survey Data Release 7 and a modified catalogue from Brinchmann+2003 and Wilman+2010. This volume complete sample has an r-band absolute magnitude limit of $M_r\leq-20$ and a redshift limit of $z\leq0.08$. See Arora+2019 for details.

1 papers1 benchmarks

VFR-447

A synthetic dataset containing 447 typefaces with only one font variation for each typeface, created for visual font recognition.

1 papers3 benchmarksImages

VFR-2420

A synthetic dataset containing word images of 447 typefaces with font variations for each typeface, created for visual font recognition.

1 papers3 benchmarksImages

BPCIS (Bacterial Phase Contrast for Instance Segementation)

BPCIS is collection of 364 bacterial phase contrast images and corresponding label matrices for instance segmentation. Labels were made according to fluorescence channels where possible. Prior to manual annotation, images were automatically cropped into microcolonies and tiled into ensemble images to reduce the empty (non-cell) image regions for training and testing. Subsequent to annotation, we performed non-rigid registration of phase contrast to cell masks.

1 papers0 benchmarksImages

Audio demo files

Audio files that supplement "Treatise on Hearing: The Temporal Auditory Imaging Theory Inspired by Optics and Communication".

1 papers0 benchmarks

BCSS (Breast Cancer Semantic Segmentation)

The BCSS dataset contains over 20,000 segmentation annotations of tissue regions from breast cancer images from The Cancer Genome Atlas (TCGA). This large-scale dataset was annotated through the collaborative effort of pathologists, pathology residents, and medical students using the Digital Slide Archive. It enables the generation of highly accurate machine-learning models for tissue segmentation.

1 papers0 benchmarksBiomedical, Images, Medical

Next2You data and results dataset (Index of Supplementary Files from "Next2You: Robust Copresence Detection Based on Channel State Information")

This record serves as an index to the other dataset releases that are part of the paper "Next2You: Robust Copresence Detection Based on Channel State Information" by Mikhail Fomichev, Luis F. Abanto-Leon, Max Stiegler, Alejandro Molina, Jakob Link, Matthias Hollick, in ACM Transactions on Internet of Things (2021).

1 papers0 benchmarks

NAO (Natural Adversarial Object)

Natural Adversarial Objects (NAO) is a new dataset to evaluate the robustness of object detection models. NAO contains 7,934 images and 9,943 objects that are unmodified and representative of real-world scenarios, but cause state-of-the-art detection models to misclassify with high confidence.

1 papers15 benchmarksImages

DSurVD (Distorted Surveillance Video Database)

A large-scale dataset, namely Distorted Surveillance Video Database (DSurVD), which can be downloaded from the link: https://sites.google.com/site/sorsyuanyuan/home/dsurvd

1 papers0 benchmarks

Scifi TV Shows (Scifi TV Show Plot Summaries & Events)

A collection of long-running (80+ episodes) science fiction TV show synopses, scraped from Fandom.com wikis. Collected Nov 2017. Each episode is considered a "story".

1 papers0 benchmarksTexts

Embrapa ADD 256 (Embrapa Apples by Drones Detection Dataset)

This is a detailed description of the dataset, a data sheet for the dataset as proposed by Gebru et al.

1 papers0 benchmarksImages

Cryptics

Official dataset of Decrypting Cryptic Crosswords: Semantically Complex Wordplay Puzzles as a Target for NLP.

1 papers0 benchmarksTexts

Multilingual Terms of Service

The first annotated corpus for multilingual analysis of potentially unfair clauses in online Terms of Service. The data set comprises a total of 100 contracts, obtained from 25 documents annotated in four different languages: English, German, Italian, and Polish. For each contract, potentially unfair clauses for the consumer are annotated, for nine different unfairness categories.

1 papers0 benchmarksTexts

Archival bundle of the data used for "Predictive Auto-scaling with OpenStack Monasca" (UCC 2021)

Follow the instructions provided in the companion repo to automatically download and decompress the archive. The following files are included:

1 papers0 benchmarks

MONK's Problems

There are three MONK's problems. The domains for all MONK's problems are the same (described below). One of the MONK's problems has noise added. For each problem, the domain has been partitioned into a train and test set.

1 papers0 benchmarks
PreviousPage 411 of 1000Next