TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Rijksmuseum Challenge 2014

Dataset used for the challenge to apply computer vision techniques on art objects (paintings, sculptures, drawings etc) from the Rijksmuseum (in Amsterdam, the Netherlands).

0 papers0 benchmarks

Robot@Home dataset

The Robot-at-Home dataset (Robot@Home) is a collection of raw and processed data from five domestic settings compiled by a mobile robot equipped with 4 RGB-D cameras and a 2D laser scanner. Its main purpose is to serve as a testbed for semantic mapping algorithms through the categorization of objects and/or rooms.

0 papers0 benchmarksImages, LiDAR, RGB-D, Videos

Robotic Instruments

Provides 8x 225-frame robotic surgical videos, captured at 2 Hz, where a trained team at Intuitive Surgical has manually labelled the different parts and types. The users are invited to test their algorithms on 8x 75-frame videos and 2x 300-frame videos which act as a test set.

0 papers0 benchmarks

Runway

A runway dataset, designing features suitable for capturing outfit appearance, collecting human judgments of outfit similarity, and learning similarity functions on the features to mimic those judgments.

0 papers0 benchmarks

SBWCE (Spanish Billion Word Corpus and Embeddings)

This resource consists of an unannotated corpus of the Spanish language of nearly 1.5 billion words, compiled from different corpora and resources from the web; and a set of word vectors (or embeddings), created from this corpus using the word2vec algorithm, provided by the gensim package. These embeddings were evaluated by translating to Spanish word2vec’s word relation test set.

0 papers0 benchmarks

ScienceIE

The shared task ScienceIE at SemEval 2017 deals with automatic extraction of keyphrases from Computer Science, Material Sciences and Physics publications, as well as extracting types of keyphrases and relations between keyphrases. PROCESS, TASK and MATERIAL form the fundamental objects in scientific works. Scientific research and practice is founded upon gaining, maintaining and understanding the body of existing scientific work in specific areas related to such fundamental objects.

0 papers0 benchmarks

SoyCultivarVein

The SoyCultivarVein dataset is a publicly available dataset, which comprises 100 categories (cultivars) with 6 samples (leaf images) in each cultivar and thus has a total number of 100×6 = 600 images (Yu et al. 2019). The leaves in the SoyCultivarVein dataset are highly similar due to the fact that they all belong to the same species, making it a new and challenging dataset for the artificial intelligence and pattern analysis research community.

0 papers0 benchmarks

Time-Lapse Hyperspectral Radiance Images

These sequences of hyperspectral radiance images have been taken from scenes undergoing natural illumination changes. In each scene, hyperspectral images were acquired at about 1-hour intervals.

0 papers0 benchmarks

TME Motorway Dataset

The “Toyota Motor Europe (TME) Motorway Dataset” is composed by 28 clips for a total of approximately 27 minutes (30000+ frames) with vehicle annotation. Annotation was semi-automatically generated using laser-scanner data. Image sequences were selected from acquisition made in North Italian motorways in December 2011. This selection includes variable traffic situations, number of lanes, road curvature, and lighting, covering most of the conditions present in the complete acquisition.

0 papers0 benchmarks

TRACT (Tweets Reporting Abuse Classification Task Corpus)

TRACT is a small scale manually annotated corpus for abuse classification problem.

0 papers0 benchmarksTexts

Transient Biometrics Nails Dataset

An extended version of an experimental dataset, called Transient Biometrics Nails Dataset (TBND), was created. TBND is composed of images of the right index finger. During acquisition the subject was instructed to lay her finger over a flat white surface and a simple point-and-shoot camera was used to acquire an image without the the use of a flash. No explicit instructions with respect to force applied were given and thus the results incorporate arbitrary force differences between users and capture sessions. Acquisition was thus done in a semi-controlled environment; apart from the white background and indirect lighting, the images present variation with respect to scale, focal plane and illumination. The dataset consists of three subsets, each one compromising the same 93 subjects, but varying on acquisition date. The first subset D01 consists of images acquired on the first acquisition day. The second subset D02 is composed of images acquired one day later. The third subset D30 was

0 papers0 benchmarks

TURBID Dataset

The TURBID, is an open image dataset that has been generated to contribute with the underwater research area. TURBID consists in a collection of five different subsets of degraded images with its respective ground-truth.

0 papers0 benchmarks

TVPR (Top-View Person Re-Identification Dataset)

The TVPR (Top View Person Re-identification) dataset stores depth frames (640x480) collected using Asus Xtion Pro Live in top-view configuration. This setup choice is primarily due to the reduction of occlusions and it has also the advantage of being privacy preserving, because faces are not recorded by the camera. The use of an RGB-D camera allows to extract anthropometric features for the recognition of people passing under the camera.

0 papers0 benchmarks

Bend the Truth

"Bend the Truth" dataset contains news in six different domains: technology, education, business, sports, politics, and entertainment. The real news included in the dataset were collected from a variety of mainstream news websites predominantly in Pakistan, India, UK, and the USA. These news channels are BBC Urdu News, CNN Urdu, Express-News, Jung News, Noway Waqat, and many other reliable news websites. The fake news included in this dataset consist of fake versions of the real news in the dataset, written by professional journalists.

0 papers0 benchmarks

Urdu Sentiment Corpus (Urdu Sentiment Corpus (v1.0): Linguistic Exploration and Visualization of Labeled Datasetfor Urdu Sentiment Analysis)

Consists of Urdu tweets for the sentiment analysis and polarity detection. The dataset is consisting of tweets, such that it casts a political shadow and presents a competitive environment between two separate political parties versus the government of Pakistan. Overall, the dataset is comprising over 17, 185 tokens with 52% records as positive, and 48% records as negative. Source: Urdu Sentiment Corpus (v1.0): Linguistic Exploration and Visualization of Labeled Dataset for Urdu Sentiment Analysis

0 papers0 benchmarks

VeRi Dataset

To facilitate the research of vehicle re-identification (Re-Id), a large-scale benchmark dateset is built for vehicle Re-Id in the real-world urban surveillance scenario, named “VeRi”. The featured properties of VeRi include:

0 papers0 benchmarks

WikiLinks

A method for automatically gathering massive amounts of naturally-occurring cross-document reference data is used to create the Wikilinks dataset comprising of 40 million mentions over 3 million entities.

0 papers0 benchmarks

Wisesight Sentiment Corpus

Social media message with sentiment label (positive, neutral, negative, question).

0 papers0 benchmarks

PCN (Pedestrian Color Naming)

Pedestrian Color Naming (PCN) is a dataset for pedestrian color naming, which contains 14,213 images, each of which hand-labeled with color label for each pixel. All images in the PCN dataset are obtained from the Market- 1501 dataset.

0 papers0 benchmarksImages

TiMoS (Tropes in Movie Synopses)

Tropes in Movie Synopses (TiMoS) is a dataset of movie tropes collected from a Wikipedia-style website, TVTropes3 with 5623 movie synopses associated with 95 most occurred tropes. The movies are diverse in genre, filming year, length, and style, making the task challenging and unable to rely on patterns from a specific domain. The tropes involve character trait, role interaction, situation, and storyline, which could be sensed by a non-expert human but remains challenging for machines that have more than 100 million parameters and pre-trained with 11,000 books and the whole Wikipedia (23.97 F1 score while a human could reach 64.87).

0 papers0 benchmarksTexts
PreviousPage 638 of 1000Next