TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

ISO17 (ISO17 - MD Trajectories of C7O2H10 with total energies and atomic forces)

Description The molecules were randomly drawn from the largest set of isomers in the QM9 dataset [1] which consists of molecules with a fixed composition of atoms (C7O2H10) arranged in different chemically valid structures. It is an extension of the ismoer MD data used in [2].

1 papers0 benchmarks

Multiple Testing and Variable Selection along Least Angle Regression's path

Data used in paper entitled "Multiple Testing and Variable Selection along Least Angle Regression's path".

1 papers0 benchmarks

HYPE (PPG and Blood Pressure from a Hypertensive Population)

HYPE Dataset - Version 1.0.0

1 papers0 benchmarksMedical

GLIB: image dataset

data/images:

1 papers0 benchmarks

Narvik Road Dataset (DIT4BEARs Smart Road Dataset)

DIT4BEARs Internship Project (at UiT-The Arctic University of Norway) Dataset

1 papers0 benchmarksTexts

MSJudge

This is a challenging dataset from real courtrooms to predict the legal judgment in a reasonably encyclopedic manner by leveraging the genuine input of the case -- plaintiff's claims and court debate data, from which the case's facts are automatically recognized by comprehensively understanding the multi-role dialogues of the court debate, and then learnt to discriminate the claims so as to reach the final judgment through multi-task learning.

1 papers0 benchmarksTexts

Color-connectivity

Synthetic graph classification datasets with the task of recognizing the connectivity of same-colored nodes in 4 graphs of varying topology.

1 papers0 benchmarksGraphs

MovieGraphBenchmark

The dataset contains entities from IMDB, TheMovieDB and TheTVDB with goldstandard matches between the sources. Due to the licensing of IMDB we provide a script to build the IMDB part of the dataset yourself.

1 papers0 benchmarksGraphs

ValidData

This dataset contains a total of 11 variables. These are: 1. vectorprice: The value in local currency of the product 2. Exchange: The official exchange rate between USD and the local currency when data was extracted. 3. Usprice: The price of the product in USD 4. vectorsold: The number of items sold by the vendor when data was extracted. 5. vectorproduct: The name of the product sold by the vendor 6. country: The name of the country where the product was sold. 7. vectorquestions: The number of questions that the vendor received when data was extracted 8. goodfeedback: the number of positive feedbacks that the vendor received when data was extracted 9. neutralfeedback: the number of neutral feedback (neither positive nor negative) 10. badfeedback: the number of negative feedback. 11. Trust: Just the ratio between goodfeedback divided by goodfeedback + neutralfeedback + badfeedback

1 papers0 benchmarks

VESUS (Varied Emotion in Syntactically Uniform Speech)

The Varied Emotion in Syntactically Uniform Speech (VESUS) repository is a lexically controlled database collected by the NSA lab. Here, actors read a semantically neutral script of words, phrases, and sentences with different emotional inflections. VESUS contains 252 distinct phrases, each read by 10 actors in 5 emotional states (neutral, angry, happy, sad, fearful).

1 papers0 benchmarksSpeech

Wasserstein Distances, Geodesics and Barycenters of Merge Trees

This repository contains all the ensemble datasets (along with their meta-data) used in the manuscript "Wasserstein Distances, Geodesics and Barycenters of Merge Trees".

1 papers0 benchmarks

CADNET

We introduce the CADNET dataset, which is an annotated collection of 3,317 3D Engineering models over 43 categories. Owing to the availability of large annotated datasets and also enough computational power in the form of GPUs, many deep learning-based solutions for object classification have been proposed of late, especially in the domain of images and graphical models. Nevertheless, very few solutions have been proposed for the task of functional classification of CAD models. Hence, for this research, CAD models have been collected from Engineering Shape Benchmark (ESB), National Design Repository (NDR), and augmented with newer models created using a modeling software to form a dataset - ‘CADNET’.

1 papers0 benchmarks

TFix's Code Patches Data

The dataset contains more than 100k code patch pairs extracted from open source projects on GitHub. Each pair comes with the erroneous and the fixed version of the corresponding code snippet. Instead of the whole file, the code snippets are extracted to focus on the problematic region (error line + other lines around it). For each sample, the repository name, the commit id, and the file names are provided so that one can access the complete files in case of interest.

1 papers4 benchmarksTexts

OG RGB+D

OG RGB+D is a new gait recognition database called OG RGB+D database, which breaks through the limitation of other gait databases and includes multimodal gait data of various occlusions (self-occlusion, active occlusion, and passive occlusion) by a multiple synchronous Azure Kinect DK sensors data acquisition system (multi-Kinect SDAS) that can be also applied in security situations. Because Azure Kinect DK can simultaneously collect multimodal data to support different types of gait recognition algorithms, especially enables to effectively obtain camera-centric multi-person 3D poses, and multi view is better to deal with occlusion than single-view. In particular, the OG RGB+D database provides accurate silhouettes and the optimized human 3D joints data (OJ) by fusing data collected by multi-Kinects which are more accurate in human pose representation under occlusion.

1 papers0 benchmarks

CARLE (Cellular Automata Reinforcement Learning Environment)

CARLE is a life-like cellular automata simulator and reinforcement learning environment. CARLE is flexible, capable of simulating any of the 262,144 different rules defining Life-like cellular automaton universes. CARLE is also fast and can simulate automata universes at a rate of tens of thousands of steps per second through a combination of vectorization and GPU acceleration. Finally, CARLE is simple. Compared to high-fidelity physics simulators and video games designed for human players, CARLE's two-dimensional grid world offers a discrete, deterministic, and atomic universal playground, despite its complexity.

1 papers0 benchmarksEnvironment

Undecided Voters in US Presidential Elections

This data contains the election polls for the 2004, 2008, 2012, and 2016 US presidential election by state including data on undecided voter proportions.

1 papers0 benchmarksTabular

UHCSDB (Ultrahigh Carbon Steel micrograph DataBase)

DeCost, Hecht, Francis, Webler, Picard, and Holm. UHCSDB (Ultrahigh Carbon Steel micrograph DataBase): tools for exploring large heterogeneous microstructure datasets. accepted for publication in IMMI 2017 doi: 10.1007/s40192-017-0097-0

1 papers0 benchmarksImages

AADB2021 (Cation-coordinated conformers of 20 proteinogenic amino acids with different protonation states)

We present a data set from a first-principles study of amino-methylated and acetylated (capped) dipeptides of the 20 proteinogenic amino acids – including alternative possible side chain protonation states and their interactions with selected divalent cations (Ca$^{2+}$, Mg$^{2+}$ and Ba$^{2+}$. The data covers 21,909 stationary points on the respective potential-energy surfaces in a wide relative energy range of up to 4 eV (390 kJ/mol). Relevant properties of interest, like partial charges, were derived for the conformers.

1 papers0 benchmarks

AADB2021Ontology (Ontology representation for a data set of cation-coordinated conformers of 20 proteinogenic amino acids)

This onotology is populated with the data from AADB2021 (https://dx.doi.org/10.17172/NOMAD/2021.02.10-1). Details can be found in the related article on arXiv.org: https://arxiv.org/abs/2107.08855

1 papers0 benchmarks

GR712RC LEON3 Power Model Data

Dataset Files The official dataset files are hosted at https://dx.doi.org/10.21227/1y7r-am78.

1 papers0 benchmarks
PreviousPage 400 of 1000Next