TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

LSLF (Large-scale Labeled Face)

Consists of a large number of unconstrained multi-view and partially occluded faces.

1 papers0 benchmarksImages

LSMDC-Context

The Large Scale Movie Description Challenge (LSMDC) - Context is an augmented version of the original LSMDC dataset with movie scripts as contextual text.

1 papers0 benchmarksVideos

Mafiascum

A collection of over 700 games of Mafia, in which players are randomly assigned either deceptive or non-deceptive roles and then interact via forum postings. Over 9000 documents were compiled from the dataset, which each contained all messages written by a single player in a single game. This corpus was used to construct a set of hand-picked linguistic features based on prior deception research, as well as a set of average word vectors enriched with subword information.

1 papers0 benchmarksTexts

MalayalamMixSentiment

MalayalamMixSentiment is a Sentiment Analysis Dataset for Code-Mixed Malayalam-English.

1 papers0 benchmarksTexts

MANTRA

An annotated dataset of 4869 transient and 71207 non-transient object lightcurves built from the Catalina Real Time Transient Survey.

1 papers0 benchmarks

Marmara Turkish Coreference Resolution Corpus

Describe the Marmara Turkish Coreference Corpus, which is an annotation of the whole METU-Sabanci Turkish Treebank with mentions and coreference chains.

1 papers0 benchmarksTexts

MCIC-COCO

A large-scale machine comprehension dataset (based on the COCO images and captions).

1 papers0 benchmarks

MD4K

A small-scale training set, which only contains 4K images.

1 papers0 benchmarks

Medical Case Report Corpus

Medical Case Report Corpus is a new corpus comprising annotations of medical entities in case reports, originating from PubMed Central's open access library.

1 papers0 benchmarksMedical, Texts

medisim

medisim is a collection of new large-scale medical term similarity datasets based on SNOMED-CT.

1 papers0 benchmarksTexts

Medley2K

A dataset called Medley2K that consists of 2,000 medleys and 7,712 labeled transitions.

1 papers0 benchmarksAudio

Mega-COV

Mega-COV is a billion-scale dataset from Twitter for studying COVID-19. The dataset is diverse (covers 234 countries), longitudinal (goes as back as 2007), multilingual (comes in 65 languages), and has a significant number of location-tagged tweets (~32M tweets).

1 papers0 benchmarksTexts

Metaphorics

Metaphorics is a newly introduced non-contextual skeleton action dataset. All the datasets introduced so far in the skeleton human action recognition have categories based only on verb-based actions.

1 papers0 benchmarksVideos

METU-ALET

METU-ALET is an image dataset for the detection of the tools in the wild. The dataset has annotations for tools that belongs to the categories such as farming, gardening, office, stonemasonry, vehicle, woodworking and workshop. The images in the dataset contains a total of 22,841 bounding boxes and 49 different tool categories.

1 papers0 benchmarksImages

MIDAS-KIKI

Consists of manually annotated dangerous and non-dangerous Kiki challenge videos.

1 papers0 benchmarksVideos

MK-SQuIT

An example dataset of 110,000 question/query pairs across four WikiData domains.

1 papers0 benchmarksTexts

MLS (Multiple Light Source)

The Multiple Light Source dataset (MLS) is a collection of 24 multiple object scenes each recorded under 18 multiple light source illumination scenarios. The illuminants are varying in dominant spectral colours, intensity and distance from the scene. The dataset can be used for the evaluation of computational colour constancy algorithms. Along with the images of the scenes the spectral characteristics of the camera, light sources and the objects are also provided, and each image includes pixel-by-pixel ground truth annotation of uniformly coloured object surfaces thus making this useful for benchmarking colour-based image segmentation algorithms.

1 papers0 benchmarksImages

MNIST-MIX

MNIST-MIX is a multi-language handwritten digit recognition dataset. It contains digits from 10 different languages.

1 papers0 benchmarksImages

Modern Hebrew Sentiment Dataset

Modern Hebrew Sentiment Dataset is a sentiment analysis benchmark for Hebrew, based on 12K social media comments, and provide two instances of these data: in token-based and morpheme-based settings.

1 papers0 benchmarksTexts

Mouse Reach

A large, annotated video dataset of mice performing a sequence of actions. The dataset was collected and labeled by experts for the purpose of neuroscience research.

1 papers0 benchmarksVideos
PreviousPage 372 of 1000Next