TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

NTIC Screening Dataset (Niramai Thermal Image for COVID19 Screening)

In the last two years, millions of lives have been lost due to COVID-19. Despite the vaccination programmes for a year, hospitalization rates and deaths are still high due to the new variants of COVID-19. Stringent guidelines and COVID-19 screening measures such as temperature check and mask check at all public places are helping reduce the spread of COVID-19. Visual inspections to ensure these screening measures can be taxing and erroneous. Automated inspection ensures an effective and accurate screening.

1 papers0 benchmarksImages, Videos

LARQS (An Evaluation Dataset for Chinese Codex Word Embedding Model)

Word embedding is a modern distributed word representations approach widely used in many natural language processing tasks. Converting the vocabulary in a legal document into a word embedding model facilitates subjecting legal documents to machine learning, deep learning, and other algorithms and subsequently performing the downstream tasks of natural language processing vis-à-vis, for instance, document classification, contract review, and machine translation. The most common and practical approach of accuracy evaluation with the word embedding model uses a benchmark set with linguistic rules or the relationship between words to perform analogy reasoning via algebraic calculation. This paper proposes establishing a 1,134 Legal Analogical Reasoning Questions Set (LARQS) from the 2,388 Chinese Codex corpus using five kinds of legal relations, which are then used to evaluate the accuracy of the Chinese word embedding model. Moreover, we discovered that legal relations might be ubiquitous

1 papers0 benchmarks

DGTA-Cattle (DeepGTAV-Cattle)

Object Detection data set created from the engine DeepGTAV, which is based on the video game GTAV. Part of the three data sets proposed in the paper. This data set is motivated from the Cattle dataset with almost the same classes.

1 papers0 benchmarksImages

Cattle

Cattle data set, which was introduced in a paper. We (not the authors) created a train-val-test split.

1 papers0 benchmarksImages

Example EPCIS Event Chain

This is an example data set for a hypothetical electronic products supply network.

1 papers0 benchmarksTracking

RoomEnv-v0 (The Room environment - v0)

The Room environment - v0

1 papers1 benchmarksGraphs, Texts

bladderbatch

Microarray gene expression data on 57 bladder samples from 5 batches.

1 papers0 benchmarks

RASFF

In the actual globalized world, the transportation of goods between any country is something normal. Considering that the protocols in quality and security vary from one country to another, there is a risk with the products that do not comply with the legislation of a country cross the border. In the case of edible products, the importance of avoiding this kind of situation is even higher. Since 1979, European Union members were obligated to register any risk to public health-related with the food and feed that is traded alongside the territory. This information has been registered in a portal called Rapid Alert System for Food and Feed (RASFF). The content of this paper provides a deep description of a set of records that goes from September 1979 to September 2019 both included. Each record represents an issue registered by RASFF workers containing a set of generic features that all issues have in common, and a set of features that are considered details of the issue. The nature of th

1 papers0 benchmarks

CNN Filter DB-Robust

Dataset for the Paper "Adversarial Robustness through the Lens of Convolutional Filters".

1 papers0 benchmarks

Korean UnSmile Dataset (SmilegateAI Korean UnSmile Dataset)

1.9K Korean Online Hate Speech Comments for Multilabel Classification (Annotated by Three Independent Labelers per Data)

1 papers0 benchmarksTexts

Multispectral Image Database

We present a database of multispectral images that were used to emulate the GAP camera. The images are of a wide variety of real-world materials and objects. We are making this database available to the research community. Details of the database can be found in the following publication:

1 papers0 benchmarks

CP2A dataset (CARLA Pedestrian Action Anticipation dataset)

We present a new simulated dataset for pedestrian action anticipation collected using the CARLA simulator. To generate this dataset, we place a camera sensor on the ego-vehicle in the Carla environment and set the parameters to those of the camera used to record the PIE dataset (i.e., 1920x1080, 110° FOV). Then, we compute bounding boxes for each pedestrian interacting with the ego vehicle as seen through the camera's field of view. We generated the data in two urban environments available in the CARLA simulator: Town02 and Town03.

1 papers0 benchmarksActions, Tracking

OSS for Social Good Project List (Supplementary Material for Leaving My Fingerprints: Motivations and Challenges of Contributing to OSS for Social Good)

Leaving My Fingerprints: Motivations and Challenges of Contributing to OSS for Social Good -> ICSE 2021 <-

1 papers0 benchmarks

CPSC2019 (The 2nd China Physiological Signal Challenge (CPSC 2019))

Introduction The China Physiological Signal Challenge 2019 (CPSC 2019) aims to encourage the development of algorithms for challenging QRS detection and heart rate (HR) estimation from short-term single-lead ECG recordings usually with low signal quality and/or abnormal rhythm waveforms.

1 papers0 benchmarksMedical

CPSC2020 (The 3rd China Physiological Signal Challenge 2020)

Introduction Abnormality of cardiac conduction system can induce arrhythmia. Abnormal heart rhythm can lead to other cardiac diseases and complications, and can be life-threatening 1. There are various types of arrhythmias and each type is associated with a pattern, and as such, it is possible to be identified. Arrhythmias can be classified into two major categories. The first category consists of arrhythmias formed by a single irregular heartbeat in electrocardiogram (ECG), herein called morphological arrhythmia, while another category consists of arrhythmias formed by a set of irregular heartbeats in ECG, herein called rhythmic arrhythmias 2. Dynamic electrocardiogram (DCG), like ECG Holter, provides an important way to monitor the incidences of arrhythmias in daily life, facilitating the doctors to check a total number and distribution of arrhythmias in a long time and thus to provide the required therapy to prevent further problems. The 3rd China Physiological Signal Challenge 2020

1 papers0 benchmarksMedical

CPSC2021 (The 4th China Physiological Signal Challenge 2021)

Introduction The 4th China Physiological Signal Challenge 2021 (CPSC 2021) aims to encourage the development of algorithms for searching the paroxysmal atrial fibrillation (PAF) events from dynamic ECG recordings.

1 papers0 benchmarksMedical

SSD_ID (Sub-Slot Dialogue dataset id number domain)

SSD (Sub-slot Dialog) dataset: This is the dataset for the ACL 2022 paper "A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots".

1 papers0 benchmarksTexts

SSD_NAME (Sub-Slot Dialogue dataset name domain)

SSD (Sub-slot Dialog) dataset: This is the dataset for the ACL 2022 paper "A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots".

1 papers6 benchmarksTexts

SSD_PLATE (Sub-Slot Dialogue dataset license plate number domain)

SSD (Sub-slot Dialog) dataset: This is the dataset for the ACL 2022 paper "A Slot Is Not Built in One Utterance: Spoken Language Dialogs with Sub-Slots".

1 papers0 benchmarksTexts

ChAII - Hindi and Tamil Question Answering

The dataset covers Hindi and Tamil, collected without the use of translation. It provides a realistic information-seeking task with questions written by native-speaking expert data annotators.

1 papers1 benchmarksTexts
PreviousPage 424 of 1000Next