TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Nasa Exoplanet Archive

The NASA Exoplanet Archive is an online astronomical exoplanet and stellar catalog and data service that collates and cross-correlates astronomical data and information on exoplanets and their host stars, and provides tools to work with these data. The archive is dedicated to collecting and serving important public data sets involved in the search for and characterization of extrasolar planets and their host stars. These data include stellar parameters (such as positions, magnitudes, and temperatures), exoplanet parameters (such as masses and orbital parameters) and discovery/characterization data (such as published radial velocity curves, photometric light curves, images, and spectra).

1 papers0 benchmarks

PolyU-BPCoMa (HK PolyU Backpack Colorized Mapping)

PolyU-BPCoMa: A Dataset and Benchmark Towards Mobile Colorized Mapping Using a Backpack Multisensorial System

1 papers0 benchmarks3D, Images, LiDAR

IEIs (Ion and Electron Insulators)

We would like to introduce three types of ion and electron insulators, i.e. Li-ion & electron insulators (LEIs), Na-ion & electron insulators (NEIs), and K-ion & electron insulators (KEIs), and provide a set of codes here to screen candidate materials from computational material database, Materials Project. The IEI materials are able to block the transport of multiple charge carriers (ions and electrons) and stay thermodynamically stable against specific alkali-metals. The screening workflows and usage of IEI materials in rechargeable solid-state Li/Na/K metal batteries are presented in the paper below.

1 papers0 benchmarksTabular

Hyperbard

Hyperbard is a dataset of diverse relational data representations derived from Shakespeare's plays. Our representations range from simple graphs capturing character co-occurrence in single scenes to hypergraphs encoding complex communication settings and character contributions as hyperedges with edge-specific node weights. By making multiple intuitive representations readily available for experimentation, we facilitate rigorous representation robustness checks in graph learning, graph mining, and network analysis, highlighting the advantages and drawbacks of specific representations.

1 papers0 benchmarks

Example dataset for CellCluster code

Dataset to be used with the https://github.com/MathBioCU/WSINDy_CellCluster code

1 papers0 benchmarks

FixEval

We introduce FixEval , a dataset for competitive programming bug fixing along with a comprehensive test suite and show the necessity of execution based evaluation compared to suboptimal match based evaluation metrics like BLEU, CodeBLEU, Syntax Match, Exact Match etc.

1 papers0 benchmarks

MICCAI'2015 Gland Segmentation Challenge Contest Dataset

MICCAI'2015 Gland Segmentation Challenge Contest Dataset Welcome to the challenge on gland segmentation in histology images. This challenge was held in conjuction with MICCAI 2015, Munich, Germany.

1 papers0 benchmarks

EVI

The EVI dataset is a challenging, multilingual spoken-dialogue dataset with 5,506 dialogues in English, Polish, and French. The dataset can be used to develop and benchmark conversational systems for user authentication tasks, i.e. speaker enrolment (E), speaker verification (V), speaker identification (I).

1 papers0 benchmarksDialog, Speech, Tabular, Texts

EEG and P300 database to determine the signal to noise ratio during a variety of realistic tasks

This database contains EEG and evoked potential recordings from 20 participants. This allows to assess the signal to noise ratio: - Signal: The P300 power and VEP power can be used to assess the signal power - Noise: The signal power consisting of EMG and baseline EEG during the different tasks allows to determine the noise level

1 papers0 benchmarks

RPCD (Reddit Photo Critique Dataset)

The Reddit Photo Critique Dataset (RPCD) contains tuples of image and photo critiques. RPCD consists of 74K images and 220K comments and is collected from a Reddit community used by hobbyists and professional photographers to improve their photography skills by leveraging constructive community feedback.

1 papers0 benchmarksImages, Texts

Matlab code for the article: Model-based selection of most informative diagnostic tests and test parameters

Description TBC

1 papers0 benchmarks

Traditional and Context-specific Spam Twitter

This data set is being released to support the spam and context-specific spam detection tasks on Twitter data.

1 papers1 benchmarksTexts

adVFed (Tencent Federated Advertising CVR Dataset)

Natural Vertical Partitioned CVR Dataset for Vertical Federated Learning

1 papers0 benchmarksTabular

COCO-MEBOW (Monocular Estimation of Body Orientation In the Wild)

COCO-MEBOW (Monocular Estimation of Body Orientation in the Wild) is a new large-scale dataset for orientation estimation from a single in-the-wild image. The body-orientation labels for 133380 human bodies within 55K images from the COCO dataset have been collected using an efficient and high-precision annotation pipeline. There are 127844 human instance in training set and 5536 human instance in validation set.

1 papers0 benchmarks

PRTiger

Dataset for automatic pull request title generation.

1 papers0 benchmarks

SRSD-Feynman (Easy set)

Our SRSD (Feynman) datasets are designed to discuss the performance of Symbolic Regression for Scientific Discovery. We carefully reviewed the properties of each formula and its variables in the Feynman Symbolic Regression Database to design reasonably realistic sampling range of values so that our SRSD datasets can be used for evaluating the potential of SRSD such as whether or not an SR method con (re)discover physical laws from such datasets.

1 papers0 benchmarksTables, Tabular

SRSD-Feynman (Hard set)

Our SRSD (Feynman) datasets are designed to discuss the performance of Symbolic Regression for Scientific Discovery. We carefully reviewed the properties of each formula and its variables in the Feynman Symbolic Regression Database to design reasonably realistic sampling range of values so that our SRSD datasets can be used for evaluating the potential of SRSD such as whether or not an SR method con (re)discover physical laws from such datasets.

1 papers0 benchmarksTables, Tabular

SRSD-Feynman (Medium set)

Our SRSD (Feynman) datasets are designed to discuss the performance of Symbolic Regression for Scientific Discovery. We carefully reviewed the properties of each formula and its variables in the Feynman Symbolic Regression Database to design reasonably realistic sampling range of values so that our SRSD datasets can be used for evaluating the potential of SRSD such as whether or not an SR method con (re)discover physical laws from such datasets.

1 papers0 benchmarksTables, Tabular

23 Pairs of Identical Twins Face Image Data

Description: 23 Pairs of Identical Twins Face Image Data. The collecting scenes includes indoor and outdoor scenes. The subjects are Chinese males and females. The data diversity inlcudes multiple face angles, multiple face postures, close-up of eyes, multiple light conditions and multiple age groups. This dataset can be used for tasks such as twins' face recognition.

1 papers0 benchmarksImages

50stateSimulations (50-State Redistricting Simulations)

Every decade following the Census, states and municipalities must redraw districts for Congress, state houses, city councils, and more. The goal of the 50-State Simulation Project is to enable researchers, practitioners, and the general public to use cutting-edge redistricting simulation analysis to evaluate enacted congressional districts.

1 papers0 benchmarks
PreviousPage 431 of 1000Next