TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Application of PanDict system based on EPSEIRV and SI3R models in epidemic forecasting and healthcare resource planning (LBH)

Global epidemics, like COVID-19, have substantial impacts on almost all countries in multiple aspects, such as economy, hospitalization, lifestyle, etc1, 2. COVID-19 can spread to populations worldwide due, in part, to their high contagiousness, but more importantly, because of our inability to quickly address some of the most fundamental problems of a newly emerged virus: 1) How quickly will the virus spread? Whether and under what conditions will new variants emerge? 3) How do we arrange our resources accordingly? Since previous epidemic models were incapable of addressing these three most important questions, we developed the PanDict system, which can help address all three of the most essential problems discussed above. To elaborate, our model consists of three crucial parts, each tackling one of the three above-mentioned problems: 1) predicting the spread of the new virus in each local community and calculating its R0 value using our newly devised EPSEIRV model; 2) creating and

0 papers0 benchmarks

WiSARD (Wilderness Search and Rescue Dataset)

WiSARD stands for Wilderness Search and Rescue Dataset (pronounced "wizard"). WiSARD consists of visual and thermal imagery taken from a drone flying over various wilderness environments in Washington, USA. The purpose of the WiSAR Image Dataset is to advance computer vision and deep learning research with a targeted application for wilderness search and rescue.

0 papers0 benchmarks

PJM(AEP)

PJM Hourly Energy Consumption Data PJM Interconnection LLC (PJM) is a regional transmission organization (RTO) in the United States. It is part of the Eastern Interconnection grid operating an electric transmission system serving all or parts of Delaware, Illinois, Indiana, Kentucky, Maryland, Michigan, New Jersey, North Carolina, Ohio, Pennsylvania, Tennessee, Virginia, West Virginia, and the District of Columbia.

0 papers0 benchmarksTime series

MCCSD (Mandarin Chinese Cued Speech Dataset)

This MCCS dataset is the first large-scale Mandarin Chinese Cued Speech dataset. This dataset covers 23 major categories of scenarios (e.g, communication, transportation and shoping) and 72 subcategories of scenarios (e.g, meeting, dating and introduction). It is recorded by four skilled native Mandarn Chinese Cued Speech cuers with portable cameras on the mobile phones. The Cued Speech videos are recorded with 30fps and 1280x720 format. We provide the raw Cued Speech videos, text file (with 1000 sentences) and corresponding annotations which contains two kind of data annotation. One is continuious video annotation with ELAN, the other is discrete audio annotations with Praat.

0 papers0 benchmarksActions, Audio, Speech, Videos

AndroDrift

Dataset for the paper entitled "Efficient Concept Drift Handling for Batch Android Malware Detection Models". Contains 100 monthly goodware and malware samples between january of 2012 and december of 2019. The training set used consist of samples for the full 2012 year, whereas the remaining data is used for evaluation purposes on a quaterly basis.

0 papers0 benchmarks

Noise-SF

Based on RADDLE and SNIPS , we construct Noise-SF, which includes two different perturbation settings. For single perturbations setting, we include five types of noisy utterances (character-level: \textbf{Typos}, word-level: \textbf{Speech}, and sentence-level: \textbf{Simplification}, \textbf{Verbose}, and \textbf{Paraphrase}) from RADDLE. For mixed perturbations setting, we utilize TextFlint to introduce character-level perturbation (\textbf{EntTypos}), word-level perturbation (\textbf{Subword}), and sentence-level perturbation (\textbf{AppendIrr}) and combine them to get a mixed perturbations dataset.

0 papers0 benchmarks

Big-Five Backstage

The dataset consists of 3265 text samples corresponding to the concatenation of lines spoken by fictional characters. Texts are extracted from 400 theatre plays written by 132 different authors. Overall, it contains 3419136 words in total with a mean equal to 1047.2 words per character. Text entries have binary labels representing gender of a character (Male or Female) and their five personality traits (Extraversion, Agreeableness, Openness, Neuroticism, Conscientiousness). The auxiliary part of the dataset includes author-level labels reflecting their gender, country of origin, and years of life.

0 papers0 benchmarksTexts

ELAI-Dust Storm (ELAI Dust Storm Dataset from MODIS)

Context As mentioned in the reference paper:

0 papers0 benchmarksImages

NEMO (NEMO: A Database for Emotion Analysis Using Functional Near-Infrared Spectroscopy)

We present a dataset for the analysis of human affective states using functional near-infrared spectroscopy (fNIRS). Data were recorded from thirty-one participants who engaged in two tasks. In the emotional perception task the participants passively viewed images sampled from the standard international affective picture system database, which provided ground-truth valence and arousal annotation for the stimuli. In the affective imagery task the participants actively imagined emotional scenarios followed by rating these for subjective valence and arousal. Correlates between the fNIRS signal and the valence-arousal ratings were investigated to estimate the validity of the dataset. Source-code and summaries are provided for a processing pipeline, brain activity group analysis, and estimating baseline classification performance. For classification, prediction experiments are conducted for single-trial 4-class classification of arousal and valence as well as cross-participant classificatio

0 papers0 benchmarksImages

Option Smile Volatility and Implied Probabilities Analysis

This study’s sample consists of seven corporations (Black Rock, Google, Meta, JP Morgan, Walgreens, Netflix, and Pepsico) analyzed across seven quarters beginning in 2021. The data includes the implied volatility level (annualized) for the day before, the day of, and the day following the earnings report. This information was obtained from the Bloomberg Terminal dataset BVOL. The data we read from the terminal is based on Bloomberg’s algorithm for calculating the implied volatility for different strikes. The value is the same for both calls and puts, which makes comparisons and calculations more straightforward. The dataset contains a mixture of high-growth, high-risk technology corporations that saw strong market tailwinds during the previous year and steady, high-dividend-paying equities. For a more comprehensive conclusion, we analyze the implied volatility levels across three expirations to determine the influence of each expiration. The shortest maturity spans from 1 to 4 days, wh

0 papers0 benchmarksTime series

LNDST

The Landsat collection contains 400x400 RGB pictures captured by the Landsat 8 satellite. Each image might show areas of water and other areas that are just the background. Your task is to design a classifier that identifies each pixel as either 0 (background) or 1 (water), using integer values. For training, you'll receive the RGB images in jpg format along with a corresponding mask. In this mask, a '0' indicates the background, and a '1' indicates water. However, for testing, you'll only receive the RGB image without the mask.

0 papers0 benchmarks

ManiCups

Multi-domain Image Editing Benchmark

0 papers0 benchmarksImages

IRIDIA-AF

A large paroxysmal atrial fibrillation long-term electrocardiogram monitoring database Abstract Atrial fibrillation (AF) is the most common sustained heart arrhythmia in adults. Holter monitoring, a long-term 2-lead electrocardiogram (ECG), is a key tool available to cardiologists for AF diagnosis. Machine learning (ML) and deep learning (DL) models have shown great capacity to automatically detect AF in ECG and their use as medical decision support tool is growing. Training these models rely on a few open and annotated databases. We present a new Holter monitoring database from patients with paroxysmal AF with 167 records from 152 patients, acquired from an outpatient cardiology clinic from 2006 to 2017 in Belgium. AF episodes were manually annotated and reviewed by an expert cardiologist and a specialist cardiac nurse. Records last from 19 hours up to 95 hours, divided into 24-hour files. In total, it represents 24 million seconds of annotated Holter monitoring, sampled at 200 Hz. Th

0 papers0 benchmarksMedical

Automatic Thermal Modeling (Estimation of Semiconductor Power Losses Through Automatic Thermal Modeling)

Dataset for reproducing the code of the work: Estimation of Semiconductor Power Losses Through Automatic Thermal Modeling.

0 papers0 benchmarks

SourceData-NLP (The SourceData-NLP dataset: integrating curation into scientific publishing for training large language models)

Introduction: The scientific publishing landscape is expanding rapidly, creating challenges for researchers to stay up-to-date with the evolution of the literature. Natural Language Processing (NLP) has emerged as a potent approach to automating knowledge extraction from this vast amount of publications and preprints. Tasks such as Named-Entity Recognition (NER) and Named-Entity Linking (NEL), in conjunction with context-dependent semantic interpretation, offer promising and complementary approaches to extracting structured information and revealing key concepts. Results: We present the SourceData-NLP dataset produced through the routine curation of papers during the publication process. A unique feature of this dataset is its emphasis on the annotation of bioentities in figure legends. We annotate eight classes of biomedical entities (small molecules, gene products, subcellular components, cell lines, cell types, tissues, organisms, and diseases), their role in the experimental de

0 papers0 benchmarksBiology, Biomedical, Texts

https://ogb.stanford.edu/docs/leader_linkprop/

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

0 papers0 benchmarks

RNA-Puzzles

RNA-Puzzles is a collective experiment for blind RNA structure prediction. The sequence of a solved RNA structure is confidentially communicated to participating modelling groups a couple of weeks prior to publication. The results are assessed and presented in common publications involving structuralists and modellers.

0 papers0 benchmarks

LMCQA (Legal Multiple Choice Question Answering)

This dataset contains a set of multiple-choice questions related to various legal topics. The dataset contains 20 questions covering various aspects of legal knowledge, such as the workings of the European Commission, types of legal documents, procedures in the court system, legal definitions, and European Union+United Kingdom law, among others.

0 papers0 benchmarksTexts

HEADSET (HEADSET: Human Emotion Awareness under Partial Occlusions Multimodal DataSET)

The volumetric representation of human interactions is one of the fundamental domains in the development of immersive media productions and telecommunication applications. Particularly in the context of the rapid advancement of Extended Reality (XR) applications, this volumetric data has proven to be an essential technology for future XR elaboration. In this work, we present a new multimodal database to help advance the development of immersive technologies. Our proposed database provides ethically compliant and diverse volumetric data, in particular 27 participants displaying posed facial expressions and subtle body movements while speaking, plus 11 participants wearing head-mounted displays (HMDs). The recording system consists of a volumetric capture (VoCap) studio, including 31 synchronized modules with 62 RGB cameras and 31 depth cameras. In addition to textured meshes, point clouds, and multi-view RGB-D data, we use one Lytro Illum camera for providing light field (LF) data simul

0 papers0 benchmarks3D, 3d meshes, Audio, Images, Point cloud, RGB Video, RGB-D, Videos

withoutbg100 dataset (withoutbg100 Dataset for Image Matting)

The withoutbg100 dataset consists of 100 image and alpha matte pairs. These pairs are chosen to represent a wide range of subjects and complexities, specifically crafted to enhance and test the capabilities of image background removal algorithms. The dataset includes images with complex elements such as fur and objects with varying transparency levels, providing a substantial challenge to even advanced matting techniques.

0 papers0 benchmarksImages
PreviousPage 654 of 1000Next