TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Iran's Built Heritage Binary Image Classification Dataset

Iran's Built Heritage Binary Image Classification Dataset contains approximately 10,500 CHB images gathered from four different sources:

1 papers0 benchmarksImages

ELITR Minuting Corpus

ELITR Minuting Corpus in JSON format.

1 papers0 benchmarksTexts

Switchboard Dialog Act Corpus

Switchboard Dialog Act Corpus

1 papers1 benchmarksTexts

CovidET-EXT

CovidET-EXT is a dataset that augments Zhan et al. (2022)'s abstractive dataset CovidET (in the context of the COVID-19 crisis) with extractive triggers. The result is a dataset of 1,883 Reddit posts about the COVID-19 pandemic, manually annotated with 7 fine-grained emotions (from CovidET) and their corresponding extractive triggers.

1 papers0 benchmarks

Unidecor (A unified deception corpus for cross-corpus deception detection.)

UNIDECOR is a unified corpus consolidating publicly available textual deception datasets into a common format.

1 papers0 benchmarks

MultiSum

MultiSum is a dataset for multimodal summarization (MSMO). It consists of 17 categories and 170 subcategories to encapsulate a diverse array of real-world scenarios. The dataset features:

1 papers0 benchmarksTexts, Videos

RKI and DIVI COVID-19 Data combined

This database consists of two main components; data on COVID-19 infections and data on the ICU occupancy of COVID-19 patients. The infections and the ICU occupancy are collected by the German health care departments, recorded by the Robert Koch Institute (2021) (RKI), the German federal government agency and scientific institute responsible for health reporting and disease control.

1 papers0 benchmarks

HighwayPavementCrackDetection

The image comes from the CCD camera of the highway measurement vehicle. Cracks and sealed cracks have been labeled. The form of labels is different from traditional block annotations, but uses redundant and dense annotation boxes. Some of the data is manually annotated, while others are model generated annotations that have undergone careful manual inspection.

1 papers0 benchmarksImages

mOKB6

Multilingual Open Knowledge Base Completion benchmark in 6 languages: English, Hindi, Telugu, Spanish, Portuguese, and Chinese.

1 papers0 benchmarks

probability_words_nli

This dataset tests the capabilities of language models to correctly capture the meaning of words denoting probabilities (WEP, also called verbal probabilities), e.g. words like "probably", "maybe", "surely", "impossible".

1 papers0 benchmarksTexts

Famous Keyword Twitter Replies

The "Famous Keyword Twitter Replies Dataset" is a comprehensive collection of Twitter data that focuses on popular keywords and their associated replies. This dataset contains five essential columns that provide valuable insights into the Twitter conversation dynamics:

1 papers0 benchmarksTexts

PTCGA200 (Patch TCGA in 200 microns by 512 px)

PTCGA200 is a public pathological H&E image datasets from Patch TCGA in 200 microns by 512 px.

1 papers0 benchmarksImages

OpenSpeaks Voice: Odia

OpenSpeaks Voice: Odia is a large speech dataset in the Odia language of India that is stewarded by Subhashish Panigrahi and is hosted at the O Foundation. It currently hosts over 70,000 audio files under a Universal Public Domain (CC0 1.0) Release. Of these, 66,000, hosted on Wikimedia Commons, include pronunciation of words and phrases, and the remaining 4,400 include pronunciation of sentences and are hosted on Mozilla Common Voice. The files on Wikimedia Commons were also released n 2023 as four physical media in the form of DVD-ROMs titled OpenSpeaks Voice: Odia Volume I, OpenSpeaks Voice: Odia Volume II, OpenSpeaks Voice: Balesoria-Odia Volume I, and OpenSpeaks Voice: Balesoria-Odia Volume II. The dataset uses Free/Libre and Open Source Software, primarily using web-based platforms such as Lingua Libre and Common Voice. Other tools used for this project include Kathabhidhana, developed by Panigrahi by forking the Voice Recorder for Tamil Wiktionary by Shrinivasan T, and Spell4wik

1 papers0 benchmarksAudio

Immobilized fluorescently stained zebrafish through the eXtended Field of view Light Field Microscope 2D-3D dataset

This dataset comprises three immobilized fluorescently stained zebrafish imaged through the eXtended Field of view Light Field Microscope (XLFM, also known as Fourier Light Field Microscope). The images were preprocessed with the SLNet, which extracts the sparse signals from the images (a.k.a. the neural activity).

1 papers0 benchmarks

ChinaOpen-1k

ChinaOpen is a new video dataset targeted at open-world multimodal learning, with raw data gathered from Bilibili, a popular Chinese video-sharing website. The dataset has a large webly annotated training set of videos (associated with user-generated titles and tags) and a smaller manually annotated test set of videos (with manually checked user titles / tags, manually written captions, and manual labels describing what visual objects / actions / scenes shown in the visual content).

1 papers3 benchmarksTexts, Videos

Uncorrelated Corrupted Dataset (UCD)

Uncorrelated Corrupted Dataset is an evaluation set that consists of realistic visible-infrared (V-I) corruptions allowing for models' corruption robustness evaluation. Initially proposed for multimodal person re-identification, our dataset can also be used for the evaluation of V-I cross-modal approaches. Corruptions of the visible modality are the twenty corruptions proposed by Chen & al. in the "Benchmarks for Corruption Invariant Person Re-identification" paper. Corruptions of the infrared modalities have been proposed in our paper, introducing 19 corruptions that respect the infrared modality encoding. In practice, the corruptions are applied randomly and independently to the visible and the infrared cameras, making it more suited to a not co-located camera setting.

1 papers0 benchmarks

Correlated Corrupted Dataset (CCD)

Correlated Corrupted Dataset is an evaluation set that consists of realistic visible-infrared (V-I) corruptions allowing for models' corruption robustness evaluation. Initially proposed for multimodal person re-identification, our dataset can also be used for the evaluation of V-I cross-modal approaches. Corruptions of the visible modality are the twenty corruptions proposed by Chen & al. in the "Benchmarks for Corruption Invariant Person Re-identification" paper. Corruptions of the infrared modalities have been proposed in our paper, introducing 19 corruptions that respect the infrared modality encoding. In practice, for co-located visible-infrared cameras, weather-related corruptions should, for example, affect each camera. Also, blur-related corruption would likely occur in both visible and infrared cameras. This dataset tackles this aspect by considering the eventual correlations that may occur from one modality camera to another.

1 papers0 benchmarks

Demande Dataset

Demande Dataset contains the features and probabilites of ten different functions.

1 papers0 benchmarks

FICLE (Factual Inconsistency CLassification with Explanation)

The FICLE dataset is a derivative of the FEVER dataset, which is a collection of 185,445 claims generated by modifying sentences obtained from Wikipedia. These claims were then verified without knowledge of the original sentences they were derived from. Each sample in the FEVER dataset consists of a claim sentence, a context sentence extracted from a Wikipedia URL as evidence, and a type label indicating whether the claim is supported, refuted, or lacks sufficient information.

1 papers0 benchmarksTexts

ACCT Data Repository (ACCT is a fast and accessible automatic cell counting tool using machine learning for 2D image segmentation)

This dataset is a collection of fluorescent images from mice in order to test an automatic cell counting tool that we developed. 62 images viewed from 2 or 3 different fields of views are shown. In brief, the dataset was derived from brain sections of a model for HIV-induced brain injury (HIVgp120tg), which expresses soluble gp120 envelope protein in astrocytes under the control of a modified GFAP promoter. The mice were in a mixed C57BL/6.129/SJL genetic background, and two genotypes of 9 month old male mice were selected: wild type controls (Resting, n = 3) and transgenic littermates (HIVgp120tg, Activated, n = 3). No randomization was performed. HIVgp120tg mice show among other hallmarks of human HIV neuropathology an increase in microglia numbers which indicates activation of the cells compared to non-transgenic littermate controls.

1 papers0 benchmarksBiology, Biomedical, Images, Medical
PreviousPage 463 of 1000Next