TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Sketch2aia (Mobile User Interface Sketches)

Dataset of 374 photos of hand-drawn sketches of App Inventor apps used for development of the Sketch2aia model for automatic generation of App Inventor wireframes from hand-drawn sketches.

1 papers0 benchmarksImages

An Amharic News Text classification Dataset

In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The task of collecting, labeling, annotating, and making valuable this kind of data will encourage junior researchers, schools, and machine learning practitioners to implement existing classification models in their language. In this short paper, we aim to introduce the Amharic text classification dataset that consists of more than 50k news articles that were categorized into 6 classes. This dataset is made available with easy baseline performances to encourage studies and better performance experiments.

1 papers2 benchmarksTexts

THEOStereo

THEOStereo is a dataset providing synthetic stereo image pairs and their corresponding scene depth and will be published along with 1. All images follow the omnidirectional camera model. In total, there are 31,250 omnidirectional images pairs. The training set contains 25,000 image pairs. For validation and testing there are 3,125 image pairs, respectively. For each pair, there is a ground truth depth map describing the pixel-wise distance of the object along the left camera's z-axis. The virtual omnidirectional cameras exhibit a FOV of 180 degrees and can be described using Kannala's camera model 2. The distortion parameters are k_1 = 1 and k_2 = k_3 = k_4 = k_5 = 0. The length of the stereo camera's baseline was 0.3 AU (approx. 15 cm, not 30 cm!). Please do not forget to cite 1 if you use the dataset in your work. Thank you.

1 papers0 benchmarksRGB-D, Stereo

UCF-101 VIPriors subset

The VIriors Action Recognition Challenge uses a subset of the UCF101 action recognition dataset:

1 papers0 benchmarksVideos

Tsinghua Dogs

Tsinghua Dogs is a fine-grained classification dataset for dogs, over 65% of whose images are collected from people's real life. Each dog breed in the dataset contains at least 200 images and a maximum of 7,449 images, basically in proportion to their frequency of occurrence in China, so it significantly increases the diversity for each breed over existing dataset. Furthermore, Tsinghua Dogs annotated bounding boxes of the dog’s whole body and head in each image, which can be used for supervising the training of learning algorithms as well as testing them.

1 papers0 benchmarksImages

DSBEC (Dark solitons in BECs dataset)

The data set consists of 6257 labeled images of Bose-Einstein condensates (BECs) with and without solitonic excitations, including kink solitons and solitonic vortices. Each element of the data set contains a masked image (132x164 pixels) of 2D atomic density used to train the machine learning model used in the paper "Machine-learning enhanced dark soliton detection in Bose-Einstein condensates," (https://arxiv.org/abs/2101.05404), and a label indicating the class a given image belongs to (0 indicates no solitons, 1 indicates a single soliton, and 2 indicates other excitations). The data structure file and project description are included with the data. This data set was used to train a deep convolutional neural network to automatically recognize whether or not a lone dark soliton has been created in BECs that was then implemented within an automated soliton detection and positioning system (see https://arxiv.org/abs/2101.05404 for details).

1 papers0 benchmarks

VESSEL12 (VESsel SEgmentation in the Lung 2012)

1 papers0 benchmarksImages

ConScenD

The ConScenD dataset consists of over 340 scenarios extracted from the naturalistic highway dataset highD. This scenarios can be used to test for the introduction of Level 3 Automated Lane Keeping Systems according to the UNECE R157 ALKS Regulation.

1 papers0 benchmarks

DODa (Darija Open Dataset)

Darija Open Dataset (DODa) is an open-source project for the Moroccan dialect. With more than 10,000 entries DODa is arguably the largest open-source collaborative project for Darija-English translation built for Natural Language Processing purposes. In fact, besides semantic categorization, DODa also adopts a syntactic one, presents words under different spellings, offers verb-to-noun and masculine-to-feminine correspondences, contains the conjugation of hundreds of verbs in different tenses, and many other subsets to help researchers better understand and study Moroccan dialect.

1 papers0 benchmarksTexts

LeT-Mi (Levantine Twitter dataset for Misogynistic language)

Levantine Twitter dataset for Misogynistic language (LeT-Mi) is an Arabic Levantine Twitter dataset for misogynistic language to be the first benchmark dataset for Arabic misogyny.

1 papers0 benchmarksTexts

HW-NAS-Bench

HW-NAS-Bench is a dataset for HardWare-aware Neural Architecture Search (HW-NAS). It is the first dataset for HW-NAS research aiming to democratize HW-NAS research to non-hardware experts and facilitate a unified benchmark for HW-NAS to make HW-NAS research more reproducible and accessible, covering two SOTA NAS search spaces including NAS-Bench-201 and FBNet

1 papers0 benchmarks

Penson et al.'s dataset derived from the MSK-IMPACT dataset

The dataset is derived from the MSK-IMPACT dataset designed and published by Zehir using the code published by Penson et al.. The derivation process is described in Development of Genome-Derived Tumor Type Prediction to Inform Clinical Cancer Care.

1 papers0 benchmarks

Cross-Linguistic Polysemies (Data from: Using network approaches to enhance the analysis of cross-linguistic polysemies)

Data from: Using network approaches to enhance the analysis of cross-linguistic polysemies

1 papers0 benchmarks

Autoencoder Paraphrase Dataset (AEPD)

This is a benchmark for neural paraphrase detection, to differentiate between original and machine-generated content.

1 papers0 benchmarksTexts

VCAS-Motion (Video Class Agnostic Segmentation Benchmark)

Video class agnostic segmentation (VCAS) is the task of segmenting objects without regards to its semantics combining appearance, motion and geometry from monocular video sequences. The main motivation behind this is to account for unknown objects in the scene and to act as a redundant signal along with the segmentation of known classes for better safety as shown in the following Figure.

1 papers0 benchmarksVideos

TREC-05 (TREC 2005 Spam Public Corpora)

1 papers0 benchmarks

USB (Universal-Scale Object Detection Benchmark)

The Universal-Scale object detection Benchmark (USB) is a benchmark for object detection that has variations in object scales and image domains by incorporating COCO with the recently proposed Waymo Open Dataset and Manga109-s dataset. To enable fair comparison, USB establishes different protocols by defining multiple thresholds for training epochs and evaluation image resolutions.

1 papers0 benchmarksImages

BookingDataChallenge (Booking.com Data Challenge)

The dataset contains anonymised hotel checkins. The dataset contains train and test parts, the in the test part city of the last checkin is masked. The goal is to predict this masked checkin.

1 papers0 benchmarks

UJIIndoorLoc

The UJIIndoorLoc is a Multi-Building Multi-Floor indoor localization database to test Indoor Positioning System that rely on WLAN/WiFi fingerprint.

1 papers0 benchmarks

Cell-200

Cell-200 is a a dataset of synthetic fluorescence microscopy images with cell populations generated by SIMCEP. The Cell-200 dataset consists of 200,000 $64\times 64$ grayscale images. The number of cells per image ranges from 1 to 200 and there are 1,000 images for each cell count. However, only a subset of Cell-200 with only odd cell counts and 10 images per count (1,000 training images in total) is used for the GAN training.

1 papers0 benchmarks
PreviousPage 385 of 1000Next