TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Machine Prarphrase Corpus (MPC)

This dataset is used to train and evaluate models for the detection of machine-paraphrased text.

1 papers0 benchmarksTexts

P-OCT (Peripapillary OCT Images)

The entire dataset consists of 61 different subjects, for each of which 12 radial OCT B-scans are collected at the Ophthalmology Department of Shanghai General Hospital by using DRI OCT-1 Atlantis (Topcon Corporation, Tokyo, Japan). The image size is 1024 × 992 pixels, corresponding to a field of view of 20.48 mm × 7.94 mm. For each subject, 2 radial OCT B-scans were randomly selected to ensure mutual exclusion. Two graders annotated these images manually through ITK-SNAP software into the optic disc and nine retinal layers under the supervision of a glaucoma specialist. For more details, please refer to our paper.

1 papers0 benchmarks

SBCoseg (SBCoseg Dataset)

The SBCoseg dataset includes 889 groups of images and each group consists of 18 images with a common object, leading to 16002 images in total. The whole dataset is divided into five subsets: with ECFB, with TR, with MH, with SD, and Normal (normal data). The five subsets contain 193, 251, 82, 83, and 280 image groups, respectively. Each original image is in JPG format with a pixel size of 360 ×360, and each ground-truth image is in PNG format.

1 papers2 benchmarksImages

Twitter Abusive Context

This dataset for abusive content detection in Twitter consists of two sets of annotations for the same set of tweets, one where the human annotators had access to the tweet's content and one where they didn't know the context.

1 papers0 benchmarksTexts

TCR-pMHC (10x Genomics T cell receptor peptide-MHC pairs)

10x Genomics dataset of sequenced TCRs barcoded by a panel of pMHCs (arranged on a dextramer)

1 papers0 benchmarks

TCR-CMV (T cell repertoires labelled by CMV serostatus)

Adaptive Biotechnologies' dataset of sequenced T cell repertoires labelled by patient age, HLA type, and CMV serostatus

1 papers0 benchmarks

ArtDL

ArtDL is a novel painting data set for iconography classification composed of images collected from online sources. Most of the paintings are from the Renaissance period and depict scenes or characters of Christian art. The data set is annotated with classes representing specific characters belonging to the Iconclass classification system.

1 papers2 benchmarksImages

fGn Traffic Traces

fGn series used in the article to develop the simulations.

1 papers0 benchmarks

Dry Bean Dataset

Seven different types of dry beans were used in this research, taking into account the features such as form, shape, type, and structure by the market situation. A computer vision system was developed to distinguish seven different registered varieties of dry beans with similar features in order to obtain uniform seed classification. For the classification model, images of 13,611 grains of 7 different registered dry beans were taken with a high-resolution camera. Bean images obtained by computer vision system were subjected to segmentation and feature extraction stages, and a total of 16 features; 12 dimensions and 4 shape forms, were obtained from the grains.

1 papers0 benchmarksImages

RUSS Dataset

RUSS (Rapid Universal Support Service) is a dataset that consists of a collection of 741 real-world step-by-step natural language instructions (raw and annotated) from the open web, and for each: its corresponding webpage DOM, ground-truth ThingTalk, and ground-truth actions.

1 papers0 benchmarksTexts

MSRB (Marine Snow Removal Benchmarking)

MSRB is a benchmarking dataset for marine snow removal of underwater images. Marine snow is one of the main degradation sources of underwater images that are caused by small particles, e.g., organic matter and sand, between the underwater scene and photosensors. The dataset consists of large-scale pairs of ground-truth and degraded images to calculate objective qualities for marine snow removal and to train a deep neural network. We propose two marine snow removal tasks using the dataset and show the first benchmarking results of marine snow removal.

1 papers0 benchmarksImages

U.S. Broadband Coverage

The U.S. Broadband Coverage data set is a publicly available dataset that reports broadband coverage percentages at a zip code-level. The authors have used differential privacy to guarantee that the privacy of individual households is preserved. The data set also contains error ranges estimates, providing information on the expected error introduced by differential privacy per zip code.

1 papers0 benchmarks

Win-Fail Action Understanding

First of its kind paired win-fail action understanding dataset with samples from the following domains: “General Stunts,” “Internet Wins-Fails,” “Trick Shots,” & “Party Games.” The task is to identify successful and failed attempts at various activities. Unlike existing action recognition datasets, intra-class variation is high making the task challenging, yet feasible.

1 papers3 benchmarksVideos

Multimodal PISA (Multimodal Piano Skills Assessment)

Dataset for multimodal skills assessment focusing on assessing piano player’s skill level. Annotations include player's skills level, and song difficulty level. Bounding box annotations around pianists' hands are also provided.

1 papers5 benchmarksAudio, Videos

AMT Objects

AMT Objects is a large dataset of object centric videos suitable for training and benchmarking models for generating 3D models of objects from a small number of photos of the objects. The dataset consists of multiple views of a large collection of object instances.

1 papers0 benchmarks3D, Videos

RepLab 2013

RepLab 2013 dataset uses Twitter data in English and Spanish (more than 142,000 tweets). The balance between both languages depends on the availability of data for each of the entities included in the dataset. The corpus consists of a collection of tweets referring to a selected set of 61 entities from four domains: automotive, banking, universities and music/artists. The domain selection was done to offer a variety of scenarios for reputation studies.

1 papers0 benchmarksTexts

Omiverse Object dataset

Omiverse Object is a large-scale synthetic dataset of 60,000 images including both transparent and opaque objects in different scenes. It is used for depth completion of transparent objects from a single RGB-D view.

1 papers0 benchmarksRGB-D

Auto-KWS

Auto-KWS is a dataset for customized keyword spotting, the task of detecting spoken keywords. The dataset closely resembles real world scenarios, as each recorder is assigned with an unique wake-up word and can choose their recording environment and familiar dialect freely.

1 papers0 benchmarksSpeech

Time Series Prediction Benchmarks

1 papers1 benchmarks

Criteo Attribution Modeling Dataset

Content of this dataset This dataset includes following files:

1 papers0 benchmarks
PreviousPage 386 of 1000Next