TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

MC_GRID (Multi_Channel_Grid)

Here we release the dataset (Multi_Channel_Grid, abbreviated as MC_Grid) used in our paper LIMUSE: LIGHTWEIGHT MULTI-MODAL SPEAKER EXTRACTION.

1 papers0 benchmarksAudio, Speech, Videos

Korean Hate Speech Evaluation Datasets

APEACH is the first crowd-generated Korean evaluation dataset for hate speech detection. Sentences of the dataset are created by anonymous participants using an online crowdsourcing platform DeepNatural AI.

1 papers0 benchmarksTexts

Casino Reviews (Online reviews of North American Casinos from Google Reviews)

This dataset contain online reviews gathered from google reviews written by north american casino users. explain motivations and summary of its content. Can be used to study user experience and relative research directions such as cultural impacts on latency of aspects, domain importance, sentiment analysis, opinion mining, aspect-based sentiment analysis, etc.

1 papers0 benchmarksTexts

WHAMR_ext

WHAMR_ext is an extension to the WHAMR corpus with larger RT60 values (between 1s and 3s)

1 papers5 benchmarksAudio, Speech

60k Stack Overflow Questions (60k Stack Overflow Questions from 2016-2020 classified into three categories based on their quality)

The dataset contains 60,000 Stack Overflow questions from 2016-2020, classified into three categories:

1 papers1 benchmarksTexts

Tweet IDs - Academic API Experiments (The lists of Tweet IDs of the experiments)

The lists of Tweet IDs for the experiments of the article: This Sample seems to be good enough! Assessing Coverage and Temporal Reliability of Twitter's Academic API by Juergen Pfeffer, Angelina Mooseder, Luca Hammer, Oliver Stritzel, David Garcia.

1 papers0 benchmarks

OUMVLP-Pose (Multi-View Large Population Database with Pose Sequence)

The OU-ISIR Gait Database, Multi-View Large Population Database with Pose Sequence (OUMVLP-Pose) is meant to aid research efforts in the general area of developing, testing and evaluating algorithms for model-based gait recognition.

1 papers0 benchmarksGraphs

DIC-C2DH-HeLa

HeLa cells on a flat glass Dr. G. van Cappellen. Erasmus Medical Center, Rotterdam, The Netherlands

1 papers2 benchmarks

Fluo-N2DH-SIM+

Simulated nuclei of HL60 cells stained with Hoescht

1 papers2 benchmarks

Fluo-C3DL-MDA231

MDA231 human breast carcinoma cells infected with a pMSCV vector including the GFP sequence, embedded in a collagen matrix

1 papers2 benchmarks

Dataset: Relationship extraction for knowledge graph creation from biomedical literature (Gene-Disease relationships)

This is the dataset used for classifying Gene-Disease relationship types from sentences. The dataset consists of 3 files:

1 papers1 benchmarksTexts

DeePore (Deep learning for rapid characterization of porous materials)

DeePore is a deep learning workflow for rapid estimation of a wide range of porous material properties based on the binarized micro–tomography images. By combining naturally occurring porous textures we generated 17,700 semi–real 3–D micro–structures of porous geo–materials with the size of $256^3$ voxels and 30 physical properties of each sample are calculated using physical simulations on the corresponding pore network models.

1 papers0 benchmarks

NR2R (Night RAW to RGB)

To form the collection of nighttime RAW samples, we first selected a total of 150 images with the spatial resolution at 3464×5202 from the training and validation sets provided by the night image challenge. And then these RAW images are pre-processed to best produce noise-free samples using a notable CNN based denoiser. This is because nighttime imaging experiences a very challenging situation with heavy noises incurred by high ISO setting under poor illumination condition (e.g., underexposure).

1 papers0 benchmarksImages

PeekDB

Data-set from "PEEK-An LSTM Recurrent Network for Motion Classification from Sparse Data"

1 papers0 benchmarks

Councils in Action

Using Council Data Project infrastructures (https://councildataproject.org), we assemble longitudinal municipal council meeting transcript data. This initial release of the Councils in Action dataset includes over 350 meetings of the city councils of Seattle Washington and Portland Oregon, and the county council of King County Washington.

1 papers0 benchmarksTexts

MTic (MADHAVLab Tic)

Periodic Tic sounds (T0=1s) sampled at 16kHz with duration of nearly 10s.

1 papers0 benchmarksAudio

STEW (Simultaneous Task EEG Workload Dataset)

This dataset consists of raw EEG data from 48 subjects who participated in a multitasking workload experiment utilizing the SIMKAP multitasking test. The subjects’ brain activity at rest was also recorded before the test and is included as well. The Emotiv EPOC device, with sampling frequency of 128Hz and 14 channels was used to obtain the data, with 2.5 minutes of EEG recording for each case. Subjects were also asked to rate their perceived mental workload after each stage on a rating scale of 1 to 9 and the ratings are provided in a separate file.

1 papers0 benchmarksEEG

Age and Gender (Age and Gender Dataset)

EEG signals from 60 users have been recorded whose age range lies between 6 and 55 years. Among all, there were 25 females and 35 male users. In general, all the participants were either school children or belonged to the socioeconomic cross section of the population with no medical history. The EEG recordings were acquired from all 14 electrodes operating at a sampling rate of 128 Hz. During recording, the participants were asked to comfortably sit on the chair with clear thoughts and a relaxed state.

1 papers0 benchmarksEEG

Replication Data for: "Deciphering Bitcoin Blockchain Data by Cohort Analysis" Version 3.1

Bitcoin is a peer-to-peer electronic payment system that popularized rapidly in recent years. Usually, we need to query the complete history of bitcoin blockchain data to acquire variables of economic meaning. This becomes increasingly difficult now with over 1.6 billion historical transactions on the Bitcoin blockchain. It is thus important to query Bitcoin transaction data in a way that is more efficient and provides economic insights. We apply cohort analysis that interprets bitcoin blockchain data using methods developed for population data in social science. Specifically, we query and process the Bitcoin transaction input and output data within each daily cohort. With this, we then create datasets and visualizations for some key indicators of bitcoin transactions, including the daily lifespan distributions of accumulated spent transaction output (STXO) and the daily age distributions of accumulated unspent transaction output (UTXO). We provide a computationally feasible approach t

1 papers0 benchmarksTabular

USC-GRAD-STDdb (Small Target Detection database)

USC-GRAD-STDdb comprises 115 video segments containing more than 25,000 annotated frames of HD 720p resolution (≈1280x720) with small objects of interest from 16 (≈4x4) to 256 (≈16x16) as pixel area. The length of the videos changes from 150 up to 500 frames. The size of every object is determined through the bounding box, so that a good annotation is of utmost importance for reliable performance metrics. As it may seem obvious, the smaller the object, the harder the annotation. The annotation has been carried out with the ViTBAT tool, adjusting the boxes as much as possible to the objects of interest in each video frame. In total, more than 56,000 ground truth labels have been generated.

1 papers10 benchmarksVideos
PreviousPage 425 of 1000Next