TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Congolese Swahili – French parallel text corpora

French sentences are sourced from Tatoeba repository and then translated into Congolese Swahili.

1 papers0 benchmarksTexts

language-modeling-recommendation

This is the Big-Bench version of our language-based movie recommendation dataset

1 papers1 benchmarksTexts

Binomial and toric ideal data (Binomial and toric ideal data for learning a performance metric of Buchberger's algorithm)

This data set consists of randomly generated binomial and toric ideals. It was used for predicting a certain complexity measure of Buchberger's algorithm for toric and binomial ideals in small number of variables.

1 papers0 benchmarks

BreastDICOM4 ([MIMBCD-UI] UTA4: Medical Imaging DICOM Files Dataset)

Several datasets are fostering innovation in higher-level functions for everyone, everywhere. By providing this repository, we hope to encourage the research community to focus on hard problems. In this repository, we present our medical imaging DICOM files of patients from our User Tests and Analysis 4 (UTA4) study. Here, we provide a dataset of the used medical images during the UTA4 tasks. This repository and respective dataset should be paired with the dataset-uta4-rates repository dataset. Work and results are published on a top Human-Computer Interaction (HCI) conference named AVI 2020 (page). Results were analyzed and interpreted on our Statistical Analysis charts. The user tests were made in clinical institutions, where clinicians diagnose several patients for a Single-Modality vs Multi-Modality comparison. For example, in these tests, we used both prototype-single-modality and prototype-multi-modality repositories for the comparison. On the same hand, the hereby dataset repres

1 papers2 benchmarksBiomedical, MRI, Medical

PSB2 (The Second Program Synthesis Benchmark Suite)

https://arxiv.org/abs/2106.06086

1 papers0 benchmarks

BreastRates4 ([MIMBCD-UI] UTA4: Rates Dataset)

Several datasets are fostering innovation in higher-level functions for everyone, everywhere. By providing this repository, we hope to encourage the research community to focus on hard problems. In this repository, we present our severity rates (BIRADS) of clinicians while diagnosing several patients from our User Tests and Analysis 4 (UTA4) study. Here, we provide a dataset for the measurements of severity rates (BIRADS) concerning the patient diagnostic. Work and results are published on a top Human-Computer Interaction (HCI) conference named AVI 2020 (page). Results were analyzed and interpreted from our Statistical Analysis charts. The user tests were made in clinical institutions, where clinicians diagnose several patients for a Single-Modality vs Multi-Modality comparison. For example, in these tests, we used both prototype-single-modality and prototype-multi-modality repositories for the comparison. On the same hand, the hereby dataset represents the pieces of information of bot

1 papers0 benchmarksBiomedical, Images, Medical, Tabular

Riposte! (Riposte! A Large Corpus of Counter-Arguments)

From the Riposte! A Large Corpus of Counter-Arguments abstract:

1 papers0 benchmarks

RESPIRATORY AND DRUG ACTUATION DATASET

Asthma is a common, usually long-term respiratory disease with negative impact on society and the economy worldwide. Treatment involves using medical devices (inhalers) that distribute medicationto the airways, and its efficiency depends on the precision of the inhalation technique. Health monitoring systems equipped with sensors and embedded with sound signal detection enable the recognition of drug actuation and could be powerful tools for reliable audio content analysis. The RDA Suite includes a set of tools for audio processing, feature extraction and classification and is provided along with a dataset consisting of respiratory and drug actuation sounds. The classification models in RDA are implemented based on conventional and advanced machine learning and deep network architectures. This study provides a comparative evaluation of the implemented approaches, examines potential improvements and discusses challenges and future tendencies. The central aim of this research is to ident

1 papers0 benchmarksAudio

BankNote-Net

Millions of people around the world have low or no vision. Assistive software applications have been developed for a variety of day-to-day tasks, including currency recognition. To aid with this task, we present BankNote-Net, an open dataset for assistive currency recognition. The dataset consists of a total of 24,816 embeddings of banknote images captured in a variety of assistive scenarios, spanning 17 currencies and 112 denominations. These compliant embeddings were learned using supervised contrastive learning and a MobileNetV2 architecture, and they can be used to train and test specialized downstream models for any currency, including those not covered by our dataset or for which only a few real images per denomination are available (few-shot learning). We deploy a variation of this model for public use in the last version of the Seeing AI app developed by Microsoft, which has over a 100 thousand monthly active users.

1 papers0 benchmarks

OrdinalDataset (Ordinal Encoding Data set)

It includes 10 data sets that consists of both raw data set and encoded data set where it is encoded through BERT-Sort Encoder with MLM initialization of .

1 papers1 benchmarksTexts

ICM (Image-Caption Matching Dataset)

ICM is curated for the image-text matching task. Each image has a corresponding caption text, which describes the image in detail. We first use CTR to select the most relevant pairs. Then, human annotators manually perform a 2nd round manual correction, obtaining 400,000 image-text pairs, including 200,000 positive cases and 200,000 negative cases. We keep the ratio of positive and negative pairs consistent in each of the train/val/test sets.

1 papers0 benchmarks

IQM (Image-Query Matching Dataset)

IQM is curated for the image-text matching task. Each image has a corresponding search query. We first use CTR to select the most relevant pairs. In this dataset, we randomly select image-query pairs in the candidate set after performing the cleaning process, obtaining 400,000 image-text pairs, including 200,000 positive cases and 200,000 negative cases. We keep the ratio of positive and negative pairs consistent in each of the train/val/test sets.

1 papers0 benchmarks

ICR (Image-Caption Retrieval Dataset)

In this dataset, we collect 200,000 image-text pairs. Each image has a corresponding caption text, which describes the image in detail. It contains two subtasks: image-to-text retrieval and text-to-image retrieval tasks.

1 papers0 benchmarks

IQR (Image-Query Retrieval Dataset)

IQR is proposed for the image-text retrieval task. We use 200,000 queries and the corresponding images as the annotated image-query pairs.

1 papers0 benchmarks

Mars Sample Localization

It contains grayscale mono and stereo images (NavCam and LocCam) from laboratory tests performed by a prototype rover on a martian-like testbed. The dataset can be used for artificial sample-tube detection and pose estimation. It also contains synthetic color images of the sample tube on a martian scenario created with Unreal Engine.

1 papers0 benchmarksImages, Stereo

Data for paper "Zone extrapolations in parametric timed automata"

Contains the current version of IMITATOR, all models and necessary scripts to reproduce all experiments on the benchmarks set.

1 papers0 benchmarks

Capture-24

This dataset contains Axivity AX3 wrist-worn activity tracker data that were collected from 151 participants in 2014-2016 around the Oxfordshire area. Participants were asked to wear the device in daily living for a period of roughly 24 hours, amounting to a total of almost 4,000 hours. Vicon Autograph wearable cameras and Whitehall II sleep diaries were used to obtain the ground truth activities performed during the period (e.g. sitting watching TV, walking the dog, washing dishes, sleeping), resulting in more than 2,500 hours of labelled data. Accompanying code to analyse this data is available at https://github.com/activityMonitoring/capture24. The following papers describe the data collection protocol in full: i.) Gershuny J, Harms T, Doherty A, Thomas E, Milton K, Kelly P, Foster C (2020) Testing self-report time-use diaries against objective instruments in real time. Sociological Methodology doi: 10.1177/0081175019884591; ii.) Willetts M, Hollowell S, Aslett L, Holmes C, Doherty

1 papers0 benchmarksActions, Tracking

Replication Data for: The elastic origins of tail asymmetry

Dataset and Stata codes for replicating Tables 1, 3 and Figures 1-4.

1 papers0 benchmarks

Parkinson Speech Dataset with Multiple Types of Sound Recordings Data Set

The PD database consists of training and test files. The training data belongs to 20 PWP (6 female, 14 male) and 20 healthy individuals (10 female, 10 male) who appealed at the Department of Neurology in Cerrahpasa Faculty of Medicine, Istanbul University. From all subjects, multiple types of sound recordings (26 voice samples including sustained vowels, numbers, words and short sentences) are taken. A group of 26 linear and time'“frequency based features are extracted from each voice sample. UPDRS ((Unified Parkinson's Disease Rating Scale) score of each patient which is determined by expert physician is also available in this dataset. Therefore, this dataset can also be used for regression.

1 papers0 benchmarks

CellTypeGraph Benchmark

Classifying all cells in an organ is a relevant and difficult problem from plant developmental biology. We here abstract the problem into a new benchmark for node classification in a geo-referenced graph. Solving it requires learning the spatial layout of the organ including symmetries. To allow the convenient testing of new geometrical learning methods, the benchmark of Arabidopsis thaliana ovules is made available as a PyTorch data loader, along with a large number of precomputed features.

1 papers2 benchmarksGraphs
PreviousPage 430 of 1000Next