TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Synthetic COVID-19 Chest X-ray

The Synthetic COVID-19 Chest X-ray Dataset consists of 21,295 synthetic COVID-19 chest X-ray images to be used for computer-aided diagnosis. These images, generated via an unsupervised domain adaptation approach, are of high quality.

1 papers0 benchmarksMedical

zbMATH Open dataset 2021

zbMATH Open contains over 4 million bibliographic entries with reviews or abstracts drawn from more than 3.000 journals and book series and more than 190.000 books.

1 papers0 benchmarks

Photozilla

Photozilla is a large-scale dataset which includes over 990k images belonging to 10 different photographic styles. The dataset can be used to train classification models to automatically classify the images into the relevant style.

1 papers0 benchmarksImages

WikiPII

WikiPII, an automatically labeled dataset composed of Wikipedia biography pages, annotated for personal information extraction.

1 papers0 benchmarksTexts

InFashAI (Inclusive Fashion AI)

AI algorithms, and in particular Machine Learning (ML) algorithms, learn from data tasks that have been traditionally done by humans such as: image classification, facial recognition, linguistic translation etc. To have a good generalization capability, AI algorithms must learn from sufficiently representative data, which is unfortunately not often the case. This results in a hyper-specialization of AI and its inability to perform well on new data whose distribution is too far from the one of the training set. It raises ethical questions which will undoubtedly have direct or indirect consequences on society. However, and despite biases they can entail, AI technologies are revolutionizing virtually every industry, and are forcing players in those industries to reinvent their businesses.

1 papers0 benchmarks

Russian Event2Mind

The work provides a comprehensive overview of the corpus for the Russian language for the commonsense inference task. Namely, we construct event phrases, which cover a wide range of everyday situations with labelled intents and reactions of the event main participant and emotions of other people involved.

1 papers1 benchmarksTexts

synthetic_dataset.h5

The synethetic dataset (10000 pairs of images and region, 2.95GB) is shared with the code (hdf5 dataset format).

1 papers0 benchmarksImages, Medical

Fast Linking Numbers of Loopy Structures Dataset

Copyright (C) 2021 Ante Qu antequ@cs.stanford.edu.

1 papers0 benchmarks3D, 3d meshes

GPLA-12

GPLA-12 is a new acoustic leakage dataset of gas pipelines involving 12 categories over 684 training/testing acoustic signals. The acoustic leakage signals were collected on the basis of an intact gas pipe system with external artificial leakages, and then preprocessed with structured tailoring which are turned into GPLA-12. GPLA-12 dedicates to serve as a feature learning dataset for time-series tasks and classifications.

1 papers0 benchmarksAudio

riboflavin

The dataset contains 71 samples with (normalized) expression data for 4,088 genes. The response variable is the riboflavin production rate in Bacilluss subtilis. It may be used to construct a graphical model.

1 papers0 benchmarks

Vāksañcayaḥ (Sanskrit Speech Corpus by IIT Bombay)

This Sanskrit speech corpus has more than 78 hours of audio data and contains recordings of 45,953 sentences with a sampling rate of 22KHz. The content is mainly readings of texts spanning over various Śāstras of Saṃskṛtam literature and also includes contemporary stories, radio program, extempore discourse, etc.

1 papers0 benchmarksSpeech, Texts

Amharic Error Corpus

Amharic Error Corpus is a manually annotated spelling error corpus for Amharic, lingua franca in Ethiopia. The corpus is designed to be used for the evaluation of spelling error detection and correction. The misspellings are tagged as non-word and real-word errors. In addition, the contextual information available in the corpus makes it useful in dealing with both types of spelling errors.

1 papers0 benchmarksTexts

Extended YouTube Faces (E-YTF)

The proposed Extended-YouTube Faces (E-YTF) is an extension of the famous YouTube Faces (YTF) dataset and is specifically designed to further push the challenges of face recognition by addressing the problem of open-set face identification from heterogeneous data i.e. still images vs video.

1 papers0 benchmarksImages, Videos

SSL (Small Size League)

This is a dataset to benchmark real-time embedded object detection models for RoboCup SSL (Small Size League).

1 papers0 benchmarksImages

FilmStills

FilmStills is a dataset of stills taken from a variety of films and TV shows, each concatenated with a color-compressed (with a factor of 2.667) version of itself.

1 papers0 benchmarks

LCO CR Dataset (Las Cumbres Observatory Cosmic Ray Dataset)

Cosmic rays in the LCO CR dataset are labeled accurately and consistently across many diverse observations from various instruments. To the best of our knowledge, this is the largest dataset of its kind. It consists of over 4,500 scientific images from Las Cumbres Observatory global telescope network's 23 instruments. Each sample in our dataset is a multi-extension FITS file, including three images, three corresponding CR masks, and three ignore masks.

1 papers0 benchmarksImages

Message Content Rephrasing

We introduce a new task of rephrasing for amore natural virtual assistant. Currently, vir-tual assistants work in the paradigm of intent-slot tagging and the slot values are directlypassed as-is to the execution engine. However,this setup fails in some scenarios such as mes-saging when the query given by the user needsto be changed before repeating it or sending itto another user. For example, for queries like‘ask my wife if she can pick up the kids’ or ‘re-mind me to take my pills’, we need to rephrasethe content to ‘can you pick up the kids’ and‘take your pills’. In this paper, we study theproblem of rephrasing with messaging as ause case and release a dataset of 3000 pairs oforiginal query and rephrased query. We showthat BART, a pre-trained transformers-basedmasked language model with auto-regressivedecoding, is a strong baseline for the task, andshow improvements by adding a copy-pointerand copy loss to it. We analyze different trade-offs of BART-based and LSTM-based seq2seqmodels

1 papers0 benchmarksTexts

HT Docking

HT Docking is a dataset consisting of 200 million 3D complex structures and 2D structure scores across a consistent set of 13 million ``in-stock'' molecules over 15 receptors, or binding sites, across the SARS-CoV-2 proteome. It is used to study surrogate model accuracy for protein-ligand docking.

1 papers0 benchmarks

FB15K237-Refined

FB15K237-Refined is a refined version of FB15k237 by KGRefiner.

1 papers0 benchmarks

WN18RR Refined

WN18RR Refined is a refined version of WN18RR by KGRefiner

1 papers0 benchmarks
PreviousPage 398 of 1000Next