TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

EiTB-ParCC

A large comparable corpus for Basque-Spanish was prepared, on the basis of independently-produced news by the Basque public broadcaster EiTB.

1 papers0 benchmarks

Electro-Magnetic Emanations Interception Dataset

An open data corpus of 123.610 labeled samples,

1 papers0 benchmarks

ENT-DESC

ENT-DESC involves retrieving abundant knowledge of various types of main entities from a large knowledge graph (KG), which makes the current graph-to-sequence models severely suffer from the problems of information loss and parameter explosion while generating the descriptions.

1 papers3 benchmarks

eTRIMS Image Database

The database is comprised of two datasets, the 4-Class eTRIMS Dataset with 4 annotated object classes and the 8-Class eTRIMS Dataset with 8 annotated object classes.

1 papers0 benchmarksImages

Europarl ConcoDisco Dataset

The ConcoDisco Corpus is an English-French parallel corpus with discourse relations (DRs) and discourse connectives (DCs) annotations.

1 papers0 benchmarksTexts

Event-focused Emotion Corpora for German and English

A corpus designed in analogy to the well-established English ISEAR emotion dataset.

1 papers0 benchmarks

Event-QA

Contains 1000 semantic queries and the corresponding English, German and Portuguese verbalizations for EventKG - an event-centric knowledge graph with more than 970 thousand events.

1 papers0 benchmarks

Explainable Abstract Trains

An image dataset containing simplified representations of trains. It aims to provide a platform for the application and research of algorithms for justification and explanation extraction. The dataset is accompanied by an ontology that conceptualizes and classifies the depicted trains based on their visual characteristics, allowing for a precise understanding of how each train was labeled. Each image in the dataset is annotated with multiple attributes describing the trains' features and with bounding boxes for the train elements.

1 papers0 benchmarks

Facebook Post Reactions

Collects posts (and their reactions) from Facebook pages of large supermarket chains.

1 papers0 benchmarks

Fake News Filipino Dataset

Expertly-curated benchmark dataset for fake news detection in Filipino.

1 papers0 benchmarks

FAS100K

FAS100K is a large-scale visual localization dataset. This dataset is comprised of two traverses of 238 and 130 kms respectively where the latter is a partial repeat of the former. The data was collected using stereo cameras in Australia under sunny day conditions. It covers a variety of road and environment types including urban and rural areas. The raw image data from one of the cameras streaming at 5 Hz constitutes 63,650 and 34,497 image frames for the two traverses respectively.

1 papers0 benchmarksImages

FeathersV1

The FeatherV1 dataset is a dataset for fine-grained visual classification. It contains 28,272 images of feathers categorized by 595 bird species.

1 papers0 benchmarksImages

FIGRIM (FIne-GRained Image Memorability)

This is a dataset of 9428 images, 1754 of which are target images with memorability scores. The images span 21 scene categories from the SUN database. Each scene category was chosen to contain at least 300 images of size 700x700 or greater. All images were cropped to 700x700 pixels.

1 papers0 benchmarksImages

FinChat (Finnish Chat Conversations on Everyday Topics)

Finnish chat conversation corpus and includes unscripted conversations on seven topics from people of different ages.

1 papers0 benchmarks

Fine-grained 3D Pose

A new large-scale dataset that consists of 409 fine-grained categories and 31,881 images with accurate 3D pose annotation.

1 papers0 benchmarks3D

FIW-MM (Families In Wild Multimedia)

A large-scale dataset for recognizing kinship in multimedia which extend FIW with multimedia data (i.e., video, audio, and contextual transcripts).

1 papers0 benchmarks

Fon-French Dataset

FFR Dataset is an ongoing project to collect, clean and store corpora of Fon and French sentences for machine translation from Fon-French. Fon (also called Fongbe) is an African-indigenous language spoken mostly in Benin, by about 1.7 million people. As training data is crucial to the high performance of a machine learning model, the aim of the project is to compile the largest set of training corpora for the research and design of translation and NLP models involving Fon. There are 117,029 parallel Fon-French sentences at the moment.

1 papers0 benchmarksTexts

French CASS dataset

Composed of judgments from the French Court of cassation and their corresponding summaries.

1 papers0 benchmarks

FSOCO

FSOCO is a collaborative dataset for vision-based cone detection systems in Formula Student Driverless competitions. It contains human annotated ground truth labels for both bounding boxes and instance-wise segmentation masks. The data buy-in philosophy of FSOCO asks student teams to contribute to the database first before being granted access ensuring continuous growth. By providing clear labeling guidelines and tools for a sophisticated raw image selection, new annotations are guaranteed to meet the desired quality.

1 papers0 benchmarksImages, Texts

FT Speech

FT Speech is a speech corpus created from the recorded meetings of the Danish Parliament, otherwise known as the Folketing (FT). The corpus contains over 1,800 hours of transcribed speech by a total of 434 speakers. It is significantly larger in duration, vocabulary, and amount of spontaneous speech than the existing public speech corpora for Danish, which are largely limited to read-aloud and dictation data.

1 papers0 benchmarksSpeech
PreviousPage 369 of 1000Next