TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

CHORD (CHOrus Recognition Dataset)

CHORD is the first chorus recognition dataset containing 627 songs for public use.

1 papers0 benchmarksAudio, Texts

BrazilDam Dataset

BrazilDAM is a multi sensor and multitemporal dataset that consists of multispectral images of ore tailings dams throughout Brazil. Landsat 8 and Sentinel 2 satellites that capture multispectral images over the years 2016, 2017, 2018 and 2019 were used. The dataset contains samples collected in different regions, which increases the diversity and representativeness of the characteristics of the dams.

1 papers0 benchmarksEnvironment, Images

SinGAN-Seg-polyps

SinGAN-Seg-polyps is a synthetic dataset for polyp segmentation consisting of 10,000 synthetic polyps and masks.

1 papers0 benchmarksBiomedical, Images, Medical

Antibody Watch

Antibody Watch is a dataset of text snippets extracted from over 2000 PubMed articles with annotations denoting specificity of antibodies.

1 papers0 benchmarksTexts

COVID-19 & Election

These datasets were used in the paper 'Evaluation of Thematic Coherence in Microblogs' (ACL, 2021). The data is structured as follows: each file represents a cluster of tweets which contains the tweet IDs, the journalist annotations for quality evaluation and issue identification, as well as the metric evaluation scores. Note that a set of 50 clusters, equally split between COVID-19 and Election domains, is shared between the 3 annotators and thus contains 3 labels.

1 papers0 benchmarksTexts

MultiCite

MultiCite is a dataset of 12,653 citation contexts from over 1,200 computational linguistics papers used for Citation context analysis (CCA). MultiCite contains multi-sentence, multi-label citation contexts within full paper texts.

1 papers0 benchmarksTexts

CityNet

CityNet is a multi-modal urban dataset containing data from 7 cities, each of which coming from 3 data sources, which can be used for urban computing and smart city research. The dataset consists of 3 types of raw data (city layout, taxi, meteorology) collected from 7 cities.

1 papers0 benchmarks

CrowdSpeech

CrowdSpeech is a publicly available large-scale dataset of crowdsourced audio transcriptions. It contains annotations for more than 20 hours of English speech from more than 1,000 crowd workers.

1 papers0 benchmarksSpeech

pd4ml (Physics Data for Machine Learning)

pd4ml is a collection of datasets from fundamental physics research -- including particle physics, astroparticle physics, and hadron- and nuclear physics -- for supervised machine learning studies. These datasets, containing hadronic top quarks, cosmic-ray induced air showers, phase transitions in hadronic matter, and generator-level histories, are made public to simplify future work on cross-disciplinary machine learning and transfer learning in fundamental physics.

1 papers0 benchmarksPhysics

Delaunay triangulation

Delaunay triangulation dataset for 5, 10, 15, 20 points.

1 papers0 benchmarks

ExBAN (ExBAN Corpus (Explanations for BAyesian Networks))

The ExBAN dataset: a corpus of NL explanations generated by crowd-sourced participants presented with the task of explaining simple Bayesian Network (BN) graphical representations. These explanations, in a separate collection effort, are rated for clarity and informativeness.

1 papers0 benchmarksTexts

ObMan-Ego

The ObMan-Ego is a large-scale synthetic hand dataset with egocentric scenes in which the simulated hands are provided by ObMan. The dataset is used for a hand segmentation task and its sim-to-real adaptation benchmark. Training, validation, and testing sets contain 150, 000, 6, 500, and 6, 500 images, respectively.

1 papers0 benchmarksImages

CPTC-2018

Intrusion alert dataset captured through the Collegiate Penetration Testing Competition (CPTC) 2018. Contains alerts from 6 student teams. For details, see "A Cybersecurity Dataset Derived from the National Collegiate Penetration Testing Competition" by Nathan Munaiah et al.

1 papers0 benchmarks

SURREALvols

Added information about the subject's body height and volumes of 14 individual body parts.

1 papers0 benchmarks

Steel Tube Dataset (Steel Tube Weld Defect Detection Dataset)

8 kinds of weld defects

1 papers0 benchmarksImages

Geography of Open Source Software

This dataset reports counts of active GitHub contributors (activity: 2019/2020) geolocated in early 2021. Counts are aggregated at the country level and at various regional scales. Besides countries, we report data on the EU NUTS2 level, for Brazilian, Russian, Chinese, Japanese, Indian, and US-American subnational geographies. We used a pipeline approach, attempting to infer location first from GitHub profile of a developer, then from linked Twitter accounts, then from email suffixes (country level only). Our data reports the count of developers identified by each stage of the pipeline, in case for instance one prefers to only use the GitHub account information.

1 papers0 benchmarks

PDE solutions

In this folder, you will find solutions of the following partial differential equations: - Burgers - Kortweg-de-Vries -Newell-Whitehead - Kuramoto-Sivashinsky

1 papers0 benchmarks

HumanoidRobotPose

The HumanoidRobotPose dataset is a dataset for real-time pose estimation of humanoid robots.

1 papers0 benchmarksImages

SBU-WSD-Corpus

SBU-WSD-Corpus is a corpus for Persian Word Sense Disambiguation (WSD). It is manually annotated with senses from the Persian WordNet (FarsNet) sense inventory. SBU-WSD-Corpus consists of 19 Persian documents in different domains such as Sports, Science, Arts, etc. It includes 5892 content words of Persian running text and 3371 manually sense annotated words (2073 nouns, 566 verbs, 610 adjectives, and 122 adverbs).

1 papers0 benchmarksTexts

Disaster

Disaster is a dataset that contains images collected from various sources for three different disasters: fire, water and land. Besides this, it also contains images for various damaged infrastructure due to natural or man made calamities and damaged human due to war or accidents.

1 papers0 benchmarksImages
PreviousPage 399 of 1000Next