TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

MP-3DHP: Multi-Person 3D Human Pose Dataset

Multi-Person 3D HumanPose Dataset (MP-3DHP) is a depth sensor-based dataset, which was constructed to facilitate the development of multi-person 3D pose estimation methods targeting real-world challenges. The dataset includes 177k training data and 33k validation data where both the 3D human poses and body segments are avaliable. The dataset also include 9k clean background data and 4k testing data including multi-person 3D poses.

1 papers0 benchmarks

Evaluating registrations of serial sections with distortions of the ground truths. Supplemental data

This is the supplemental data for our paper on how to benchmark registrations of serial sections with ground truths. There are three main modalities and one further, as a reference.

1 papers0 benchmarks3D, Biomedical, Medical

UTFPR-SBD3

The semantic segmentation of clothes is a challenging task due to the wide variety of clothing styles, layers and shapes. The UTFPR-SBD3 contains 4,500 images manually annotated at pixel level in 18 classes plus background. To ensure the high quality of the dataset, all images were manually annotated at the pixel level using JS Segment Annotator, 2 a free web-based image annotation tool. The raw images were carefully selected to avoid, as far as possible, classes with low number of instances.

1 papers4 benchmarksImages

FGraDA (Fine-Grained Domain Adaptation Dataset)

Previous research for adapting a general neural machine translation (NMT) model into a specific domain usually neglects the diversity in translation within the same domain, which is a core problem for domain adaptation in real- world scenarios. One representative of such challenging scenarios is to deploy a translation system for a conference with a specific topic, e.g., global warming or coronavirus, where there are usually extremely less resources due to the limited schedule. To motivate wider investigation in such a scenario, we present a real-world fine-grained domain adaptation task in machine translation (FGraDA). The FGraDA dataset consists of Chinese-English translation task for four sub-domains of information technology: autonomous vehicles, AI education, real-time networks, and smart phone. Each sub-domain is equipped with a development set and test set for evaluation pur- poses. To be closer to reality, FGraDA does not employ any in-domain bilingual training data but provide

1 papers0 benchmarks

IMDB-WIKI-SbS

IMDB-WIKI-SbS is a new large-scale dataset for evaluation pairwise comparisons, building on the success of a well-known benchmark for computer vision systems IMDB-WIKI. This dataset uses the age information offered by IMDB-WIKI as ground truth while providing a balanced distribution of ages and genders of people in photos.

1 papers0 benchmarksRanking

LIRIS human activities dataset

The LIRIS human activities dataset contains (gray/rgb/depth) videos showing people performing various activities taken from daily life (discussing, telphone calls, giving an item etc.). The dataset is fully annotated, where the annotation not only contains information on the action class but also its spatial and temporal positions in the video. It was originally shot for the ICPR-HARL 2012 competition.

1 papers0 benchmarksVideos

notebookcdg

Inspired by Wang et al. 2021, we decided to utilize the top-voted and well-documented Kaggle notebooks to construct the notebookCDGdataset

1 papers0 benchmarksTexts

IATOS Dataset

Archivos con audios de toses de personas grabadas por celular, segmentados por COVID positivo y negativo segĂșn resultado de test RT-PCR.

1 papers0 benchmarks

Orchard (A Benchmark For Measuring Systematic Generalization of Multi-Hierarchical Reasoning)

Orchard is a diagnostic dataset for systematically evaluating hierarchical reasoning in state-of-the-art neural sequence models

1 papers0 benchmarks

A dataset of neonatal EEG recordings with seizures annotations

Neonatal seizures are a common emergency in the neonatal intensive care unit (NICU). There are many questions yet to be answered regarding the temporal/spatial characteristics of seizures from different pathologies, response to medication, effects on neurodevelopment and optimal detection. This dataset contains EEG recordings from human neonates and the visual interpretation of the EEG by the human expert. Multi-channel EEG was recorded from 79 term neonates admitted to the neonatal intensive care unit (NICU) at the Helsinki University Hospital. The median recording duration was 74 minutes (IQR: 64 to 96 minutes). EEGs were annotated by three experts for the presence of seizures. An average of 460 seizures were annotated per expert in the dataset, 39 neonates had seizures by consensus and 22 were seizure free by consensus. The dataset can be used as a reference set of neonatal seizures, for the development of automated methods of seizure detection and other EEG analysis, as well as for

1 papers0 benchmarks

MIS-Check Dam (Minor Irrigation Structures- Check Dam)

Minor Irrigation Structures Check-Dam Dataset is a public dataset annotated by domain experts using images from Google static map for instance segmentation and object detection tasks.

1 papers0 benchmarksImages

Manually annotated 3-digit occupation codes from the Norwegian 1950 census

Manually annotated 3-digit occupation codes from the Norwegian full count 1950 population census.

1 papers0 benchmarks

Manually annotated 3-digit occupation code training set from the Norwegian 1950 census

The Norwegian Historical Data Centre, 2021, "Manually annotated 3-digit occupation code training set from the Norwegian 1950 census", https://doi.org/10.18710/7JWAZX, DataverseNO, V1

1 papers0 benchmarks

Sentence-level argument annotation

The dataset is based on a debate.org crawl. It is restricted to a subset of four out of the total 23 categories -- politics, society, economics and science -- and contains additional annotations. 3 human annotators familiar with linguistics segmented these documents and labeled them as being of medium or low quality, to exclude low quality documents. The annotators were then asked to indicate the beginning of each new argument and to label argumentative sentences summarizing the aspects of the post as conclusion and outside of argumentation. In this way, we obtained a ground truth of labeled arguments on a sentence level (Krippendorff's alpha=0.24 based on 20 documents and three annotators).

1 papers0 benchmarks

debatepedia

Debatepedia is a debate platform that lists arguments to a topic on one page, including subtitles, structuring the arguments into different aspects.

1 papers0 benchmarks

Student Essay

Student Essay is widely used in research on argument segmentation

1 papers0 benchmarks

Banglish

A Bilingual Dataset for Bangla and English Voice Commands

1 papers1 benchmarksAudio

washed_contract

Dataset contains about 48K contracts which are open source on Etherscan.

1 papers0 benchmarksTables

AAAC (Artificial Argument Analysis Corpus)

DeepA2 is a modular framework for deep argument analysis. DeepA2 datasets contain comprehensive logical reconstructions of informally presented arguments in short argumentative texts. This item references two two synthetic DeepA2 datasets for artificial argument analysis: AAAC01 and AAAC02.

1 papers0 benchmarks

Pre-Processed Power Grid Frequency Time Series

This repository contains ready-to-use frequency time series as well as the corresponding pre-processing scripts in python.

1 papers0 benchmarks
PreviousPage 413 of 1000Next