TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

TBCOV

TBCOV is a large-scale Twitter dataset comprising more than two billion multilingual tweets related to the COVID-19 pandemic collected worldwide over a continuous period of more than one year. Several state-of-the-art deep learning models are used to enrich the data with important attributes, including sentiment labels, named-entities (e.g., mentions of persons, organizations, locations), user types, and gender information. A geotagging method is proposed to assign country, state, county, and city information to tweets, enabling a myriad of data analysis tasks to understand real-world issues.

1 papers0 benchmarksTexts

Energy Consumption Curves of 499 Customers from Spain

Predictions of energy consumption are crucial for energy retailers to minimize deviations from energy acquired in the day-ahead market and the actual consumption of their customers. The increasing spread of smartmeters means that retailers have access to hourly consumption values of all their contracted customers in realtime. Using machine learning algorithms, these hourly values can be used to calculate predictions for the future energy consumption of the customers. The present data set allows the training and validation of AI-based prediction models.

1 papers0 benchmarksTime series

EM-POSE

Electromagnetic measurements obtained from 12 wireless sensors, paired with the corresponding ground-truth SMPL poses. Approx. 37 minutes recorded with 5 participants.

1 papers0 benchmarks

GeoMNIST (Geometric Shapes MNIST)

A simple dataset consisting of three geometric shapes (Triangle, Rectangle, Ellipsoid) of similar sizes but different orientations.

1 papers0 benchmarksImages

Gait Dataset (Human Gait Dataset)

Details about the creation of the dataset can be seen in https://arxiv.org/abs/2110.06139.

1 papers0 benchmarks3D, 6D

STR-2021

The STR-2021 dataset has 5,500 English sentence pairs manually annotated for semantic relatedness using a comparative annotation framework.

1 papers0 benchmarksTexts

MUNO21

MUNO21 is a large-scale and comprehensive dataset for the map update task. It includes time series of aerial images and map data to capture the evolution of both the physical road network and real street maps over time -- we collect NAIP aerial images at each of four years over the eight-year timespan from 2012–2019, and OSM extracts from each year during the same timespan.

1 papers0 benchmarksImages

ImageNet 50 samples per class

This ImageNet version contains only 50 training images per class while the original testing set remains unchanged. It is one of the datasets comprising the data-efficient image classification (DEIC) benchmark. It was proposed to challenge the generalization capabilities of modern image classifiers.

1 papers1 benchmarksImages

Deep Sea Treasure Pareto-Front

The dataset contains two Pareto-fronts: - The Pareto-front for the 2-objective problem - The Pareto-front for the 3-objective problem

1 papers0 benchmarksTables

POG (People On Grass)

Object detection dataset featuring people walking on grass captured aboard a UAV. This data sets include precise meta data information about altitude, viewing angle and others.

1 papers0 benchmarksImages

CCQA

CCQA is new web-scale dataset for in-domain model pre-training. CCQA is a novel QA dataset based on the Common Crawl project. Using the readily available schema.org annotation, around 130 million multilingual question-answer pairs are extracted, including about 60 million English data-points.

1 papers0 benchmarks

SHREC'16 Partial Benchmark (SHREC 2016 TRACK: PARTIAL MATCHING OF DEFORMABLE SHAPES)

Finding a correspondence between two shapes is a fundamental task in computer graphics and geometry processing with applications ranging from texture mapping to animation. A particularly challenging and widely studied setting is when shapes are allowed to undergo quasi-isometric deformations, as it happens when we consider articulated bodies in different poses. An even more interesting scenario is partial correspondence, where one is shown only a subset of the shape, and has to match it with a deformable version thereof. Partial correspondence problems arise in numerous applications that involve real data acquisition by 3D sensors, which inevitably leads to missing parts due to occlusions and partial views.

1 papers0 benchmarks

BEAMetrics

BEAMetrics (Benchmark to Evaluate Automatic Metrics) is resource to make research into new metrics for evaluation of generated language easier to evaluate. BEAMetrics users can quickly compare existing and new metrics with human judgements across a diverse set of tasks, quality dimensions (fluency vs. coherence vs. informativeness etc), and languages.

1 papers0 benchmarksTexts

NYU-VPR

NYU-VPR is a dataset for Visual place recognition (VPR) that contains more than 200,000 images over a 2km×2km area near the New York University campus, taken within the whole year of 2016.

1 papers0 benchmarksImages

Experiment-data-for-UM-S-TM

0.This is experiment data for the following article: @misc{liu2021topic, title={Topic Model Supervised by Understanding Map}, author={Gangli Liu}, year={2021}, eprint={2110.06043}, archivePrefix={arXiv}, primaryClass={cs.CL} }

1 papers0 benchmarks

NACA Airfoils

The training datasets consisting of NACA 4- and 5-digit airfoils, at different flight conditions, were generated using Javafoil.

1 papers0 benchmarks

ICASSP 2021 Acoustic Echo Cancellation Challenge

The ICASSP 2021 Acoustic Echo Cancellation Challenge is intended to stimulate research in the area of acoustic echo cancellation (AEC), which is an important part of speech enhancement and still a top issue in audio communication and conferencing systems. Many recent AEC studies report good performance on synthetic datasets where the train and test samples come from the same underlying distribution. However, the AEC performance often degrades significantly on real recordings. Also, most of the conventional objective metrics such as echo return loss enhancement (ERLE) and perceptual evaluation of speech quality (PESQ) do not correlate well with subjective speech quality tests in the presence of background noise and reverberation found in realistic environments. In this challenge, we open source two large datasets to train AEC models under both single talk and double talk scenarios. These datasets consist of recordings from more than 2,500 real audio devices and human speakers in real en

1 papers0 benchmarksAudio, Speech

INTERSPEECH 2021 Acoustic Echo Cancellation Challenge

The INTERSPEECH 2021 Acoustic Echo Cancellation Challenge is intended to stimulate research in the area of acoustic echo cancellation (AEC), which is an important part of speech enhancement and still a top issue in audio communication and conferencing systems. Many recent AEC studies report reasonable performance on synthetic datasets where the train and test samples come from the same underlying distribution. However, the AEC performance often degrades significantly on real recordings. Also, most of the conventional objective metrics such as echo return loss enhancement (ERLE) and perceptual evaluation of speech quality (PESQ) do not correlate well with subjective speech quality tests in the presence of background noise and reverberation found in realistic environments. In this challenge, we open source two large datasets to train AEC models under both single talk and double talk scenarios. These datasets consist of recordings from more than 5,000 real audio devices and human speakers

1 papers0 benchmarksAudio, Speech

FIRESTARTER 2 - dataset and notebooks

Data used in the paper "FIRESTARTER 2: Dynamic Code Generation for Processor Stress Tests", as well as notebooks to generate plots.

1 papers0 benchmarks

Small-Bench NLP

Small-Bench NLP is a benchmark for small efficient neural language models trained on a single GPU. Small-Bench NLP benchmark comprises of eight NLP tasks on the publicly available GLUE datasets and a leaderboard to track the progress of the community.

1 papers0 benchmarks
PreviousPage 409 of 1000Next