TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

CMWD (Cloud Motion Wind Dataset)

CMWD (Cloud Motion Wind Dataset) is the first cloud motion wind dataset for deep learning research. It contains 6388 adjacent grayscale image pairs for training and another 715 images pairs for testing.

1 papers0 benchmarks

TCLD (Typhoon Center Location Dataset)

TCLD (Typhoon Center Location Dataset) is a brand new typhoon center location dataset for deep learning research. It contains 1809 grayscale images for training and another 319 images for testing.

1 papers0 benchmarks

SCMD2016 (Satellite Cloudage Map Dataset)

SCMD dataset is a brand new cloudage nowcasting dataset for deep learning research. It contains 20000 grayscale image sequences for training and another 3500 image sequences for testing. You can get the SCMD2016 dataset at any time but only for scientific research. At the same time, please cite our work when you use the SCMD dataset

1 papers0 benchmarks

EviLOG (Evidential Lidar Occupancy Grid Mapping)

The dataset contains synthetic training, validation and test data for occupancy grid mapping from lidar point clouds. Additionally, real-world lidar point clouds from a test vehicle with the same lidar setup as the simulated lidar sensor is provided. Point clouds are stored as PCD files and occupancy grid maps are stored as PNG images whereas one image channel describes evidence for a free and another one describes evidence for occupied cell state.

1 papers0 benchmarksEnvironment, LiDAR, Point cloud

Viwiki-Spelling (Vietnamese Spelling Correction Dataset)

We introduce a first Vietnamese Spelling Correction dataset containing manual labelling mistakes and corresponding correct words.

1 papers0 benchmarksTexts

DiaKG

DiaKG is a high-quality Chinese dataset for Diabetes knowledge graph.

1 papers0 benchmarksTexts

MAOMaps

MAOMaps is a dataset for evaluation of Visual SLAM, RGB-D SLAM and Map Merging algorithms. It contains 40 samples with RGB and depth images, and ground truth trajectories and maps. These 40 samples are joined into 20 pairs of overlapping maps for map merging methods evaluation. The samples were collected using Matterport3D dataset and Habitat simulator.

1 papers0 benchmarksImages

Cleft

The Cleft dataset is a collection of ultrasound tongue imaging and audio data, gathered from children with cleft lip and palate by a research speech and language therapist working in a hospital environment.

1 papers0 benchmarksAudio

D-OCC (Dynamic-OneCommon Corpus)

D-OCC is a large-scale dataset of 5,617 dialogues to enable fine-grained evaluation and analysis of various dialogue systems. It is used to study common grounding in dynamic environments.

1 papers0 benchmarksTexts, Videos

Neural Closure Models - Runs

The following are all the runs used to generate figures in the paper. Every experiment solves the corresponding high- and low-fidelity model to generate the training, validation, and prediction data.

1 papers0 benchmarks

LIGHT-Quests

LIGHT-Quests is an extension of LIGHT, a large-scale crowd-sourced fantasy text-game, to generate a dataset of quests. These contain natural language motivations paired with in-game goals and human demonstrations; completing a quest might require dialogue or actions (or both).

1 papers0 benchmarksTexts

MT40K

The MT40K dataset for predicting malware threat intelligence is a collection of 40,000 triples generated from 27,354 unique entities and 34 relations. The corpus consists of approximately 1,100 de-identified plain text threat reports written between 2006-2021 and all CVE vulnerability descriptions created between 1990 to 2021. The annotated keyphrases were classified into entities derived from semantic categories defined in malware threat ontologies.

1 papers0 benchmarksTexts

Multi-template MRI mouse brain atlas (Multi-template MRI mouse brain atlas for both in vivo and ex vivo analysis)

Mouse Brain MRI atlas (both in-vivo and ex-vivo) (repository relocated from the original webpage)

1 papers0 benchmarks3D, Biomedical, Images, MRI, Medical

D3DFACS (Dynamic 3D Facial Action Coding System Database)

The D3DFACS dataset is a dynamic 3D facial expression data set based on the Facial Action Coding System. It contains Action Unit (AU) sequences from 10 people, with 519 sequences in total. The peak image of each expression sequence has been manually FACS coded by a certified expert.

1 papers0 benchmarksImages

Classic ECN AQM Fall-Back

Clickable heat-map visualizations of the experiments run to quantify the Classic ECN AQM problem and to evaluate the success of the Classic AQM Detection and Fall-back algorithm.

1 papers0 benchmarksGraphs, Images

ClueWeb09

The ClueWeb09 dataset was created to support research on information retrieval and related human language technologies. It consists of about 1 billion web pages in ten languages that were collected in January and February 2009. The dataset is used by several tracks of the TREC conference.

1 papers0 benchmarks

TRECDD (TREC Dynamic Domain)

The dataset used for TREC 2017 Dynamic Domain Track consists of two domains: Ebola and New York Times.

1 papers0 benchmarks

BugClassify

Dataset of 5,591 labeled issue tickets. Originally created by Herzig et al. in : "It’s Not a Bug, It’s a Feature: How Misclassification Impacts Bug Prediction" (paper)

1 papers0 benchmarksTexts

Webly-Reference SR Dataset

Webly-Reference SR dataset is a test dataset for evaluating Ref-SR methods. It has the following advantages:

1 papers0 benchmarks

CPNet (CorresPondenceNet)

CPNet dataset has a collection of 25 categories, 2,334 models based on ShapeNetCore, which includes 1,000+ correspondence sets with 104,861 points.

1 papers0 benchmarks3D
PreviousPage 395 of 1000Next