TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Contextualised Polyseme Word Sense Dataset v2

This is a revised and extended second version of a Contextualised Polyseme Word Sense Dataset. The dataset contains two human annotated measures of word sense similarity for polysemic target words used in contexts invoking different sense interpretations. The first set contains graded similarity judgements for highlighted target words displayed in two different contexts. The second set contains co-predication acceptability judgements for sentence constructions combining the sentence pairs from the first set.

1 papers0 benchmarksTexts

Riseholme-2021

Risholme-2021 contains >3.5K images of strawberries at various growth stages along with anomalous instances. Data collection was performed in the strawberry research farm at the Riseholme campus of the University of Lincoln in UK. For more details, please check out "Homepage" down below.

1 papers0 benchmarksImages

CCIHP (Characterized Crowd Instance-level Human Parsing)

CCIHP dataset is devoted to fine-grained description of people in the wild with localized & characterized semantic attributes. It contains 20 attribute classes and 20 characteristic classes split into 3 categories (size, pattern and color). The dataset has been introduced in this paper: Loesch, A., & Audigier, R. (2021, September). Describe me if you can! Characterized instance-level human parsing. In 2021 IEEE International Conference on Image Processing (ICIP) (pp. 2528-2532). IEEE. The annotations were made with Pixano, an opensource, smart annotation tool for computer vision applications: https://pixano.cea.fr/

1 papers0 benchmarksImages

COVID-19 Contact Tracing Survey in Israel

A survey of Israelis about their attitudes towards COVID-19 contact tracing apps

1 papers0 benchmarks

!Optimizer 2021 Data

The data used for !Optimizer 2021 competition, based on seven biological model organisms.

1 papers0 benchmarks

FooDI-ML (Food Drinks and groceries Images Multi Lingual)

Food Drinks and groceries Images Multi Lingual (FooDI-ML) is a dataset that contains over 1.5M unique images and over 9.5M store names, product names descriptions, and collection sections gathered from the Glovo application. The data made available corresponds to food, drinks and groceries products from 37 countries in Europe, the Middle East, Africa and Latin America. The dataset comprehends 33 languages, including 870K samples of languages of countries from Eastern Europe and Western Asia such as Ukrainian and Kazakh, which have been so far underrepresented in publicly available visiolinguistic datasets. The dataset also includes widely spoken languages such as Spanish and English.

1 papers0 benchmarksImages

Nico-illust (Nico-Illust)

This dataset contains over 400,000 images (illustrations) from Niconico Seiga and Niconico Shunga

1 papers0 benchmarksImages

Interference suppression techniques for OPM-based MEG: Opportunities and challenges

OPM data

1 papers0 benchmarks

ETH Kinect Dataset

This dataset contains 27 ROS bags of point clouds produced by a Kinect based the ground truth obtained from a Vicon pose capture system. These runs cover 3 environments of increasing complexity, with 3 types of motions at 3 different speeds. This dataset can be used with our ICP Mapper to track the pose of the Kinect and to explore parameters of ICP algorithms.

1 papers0 benchmarks

ETH Laser Registration Datasets

This group of datasets was recorded with the aim to test point cloud registration algorithms in specific environments and conditions. Special care is taken regarding the precision of the "ground truth" positions of the scanner, which is in the millimeter range, using a theodolite.

1 papers0 benchmarks

Labeled Retinal Optical Coherence Tomography Dataset for Classification of Normal, Drusen, and CNV Cases

This dataset consists of more than 16,000 retinal OCT B-scans from 441 cases (Normal: 120, Drusen: 160, CNV: 161) and is acquired at Noor Eye Hospital, Tehran, Iran. Images are labeled by a retinal specialist.

1 papers0 benchmarksImages, Medical

AraCovid19-SSD

AraCovid19-SSD is a manually annotated Arabic COVID-19 sarcasm and sentiment detection dataset containing 5,162 tweets.

1 papers0 benchmarksTexts

Aristo-v4 (Aristo Tuple KB Version 4)

The Aristo Tuple KB contains a collection of high-precision, domain-targeted (subject,relation,object) tuples extracted from text using a high-precision extraction pipeline, and guided by domain vocabulary constraints. The dataset was introduced by the paper Domain-Targeted, High Precision Knowledge Extraction.

1 papers4 benchmarks

HowSumm

HowSumm is a large-scale query-focused multi-document summarization dataset. It is focused on summarization of various sources to create HowTo guides. It is derived from wikiHow articles.

1 papers0 benchmarksTexts

RWD-10K (Rogue Wave Dataset-10K)

Rogue Wave Dataset-10K dataset consists of 10191 rogue wave images.

1 papers0 benchmarksImages

Odysseus (Clean and Trojan Models)

A major reason for the lack of a realistic Trojan detection method has been the unavailability of a large-scale benchmark dataset, consisting of clean and Trojan models. Here we introduce Odysseus the largest public dataset that contains over 3,000 trained clean and Trojan models based on Pytorch.

1 papers0 benchmarks

CGHD1152 (Circuit Graph Hand Drawn 1152)

1152 Images 144 Circuits 12 Drafter 48,563 Object (Symbol, Structural, Text) Annotations

1 papers0 benchmarks

A Curb Dataset

This is a dataset with curb annotations by using 3D LiDAR data and we build this dataset based on the SemanticKITTI dataset.

1 papers0 benchmarks

MOD20

MOD20 is an action recognition dataset consisting of videos collected from YouTube and our own drone. The dataset contains 2,324 videos lasting a total of 240 minutes. The actions were selected from challenging and complex scenarios, and cover multiple viewpoints, from ground-level to bird's-eye view. The substantial variation in body size, number of people, viewpoints, camera motion, and background makes this dataset challenging for action recognition. The action classes, 720×720 size un-distorted clips and multi-viewpoint video selection extend the dataset's applicability to a wider research community.

1 papers0 benchmarksVideos

NGAFID-MC

NGNGAFID-MC consists of over 7500 labeled flights, representing over 11,500 hours of per second flight data recorder readings of 23 sensor parameters.

1 papers0 benchmarks
PreviousPage 408 of 1000Next