TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

DUC 2006 (Document Understanding Conferences)

There is currently much interest and activity aimed at building powerful multi-purpose information systems. The agencies involved include DARPA, ARDA and NIST. Their programmes, for example DARPA's TIDES (Translingual Information Detection Extraction and Summarization) programme, ARDA's Advanced Question & Answering Program and NIST's TREC (Text Retrieval Conferences) programme cover a range of subprogrammes. These focus on different tasks requiring their own evaluation designs.

1 papers0 benchmarksTexts

raw gaze data

Data was collected from Tobii Fusion screen-based Eye Tracker. This study collected drivers’ gaze data by letting participants watch dashcam captured videos of driving scenes in the lab setting. Original vidoes are downloaded from https://github.com/Cogito2012/CarCrashDataset. Each video lasts 5 seconds and the frequency of the videos is 10 Hz.

1 papers0 benchmarks

Euro-PVI

The Euro-PVI dataset contains trajectories of pedestrians and bicyclists, with dense interactions with the ego-vehicle. The dataset is collected in Brussels and Leuven, Belgium. The goal of this dataset is to address the challenge of future trajectory prediction in urban environments with dense pedestrian (bicyclist) - vehicle interactions.

1 papers0 benchmarks

Knot128

Knot128 is a dataset to test knot untangling algorithms, i.e., highly-tangled configurations that can be difficult to smooth out into a canonical knot embedding. Knot128 is comprised of knots from 128 different isotopy classes; for each class, a tangled embedding, and a canonical embedding are provided.

1 papers0 benchmarks

Trefoil100

Trefoil100 is a dataset to test knot untangling algorithms, i.e., highly-tangled configurations that can be difficult to smooth out into a canonical knot embedding. Trefoil100 contains 100 tangled embeddings of the trefoil knot.

1 papers0 benchmarks

CF-mMIMO data - measurement at USC

This repo contains open-source channel measurement data for research and development purposes.

1 papers0 benchmarks

I2L-140K

Introduced by Singh, Sumeet S.. “Teaching Machines to Code: Neural Markup Generation with Visual Attention.” ArXiv abs/1802.05415 (2018): n. pag.

1 papers1 benchmarks

Im2latex-90k

Introduced by Singh, Sumeet S.. “Teaching Machines to Code: Neural Markup Generation with Visual Attention.” ArXiv abs/1802.05415 (2018): n. pag.

1 papers0 benchmarks

Interactive Media Experience Click Dataset

The dataset contains summary statistics and engagement metrics captured from users in a live, 'in-the-wild' study of an interactive TV show.

1 papers0 benchmarks

Multispectral and HD vineyard orthomosaics from central Portugal

Multispectral and HD vineyard orthomosaics from central Portugal

1 papers0 benchmarksImages

VideoRemoval4K

We provide video sequences with annotated object masks for video inpainting. The resolution is 3840 x 2160.

1 papers0 benchmarks

MedLEA (Medicinal Leaves)

The MedLEA package provides morphological and structural features of 471 medicinal plant leaves and 1099 leaf images of 31 species and 29-45 images per species.

1 papers0 benchmarks

Swedish Leaf Dataset

A dataset of images containing leaves from 15 tree classes.

1 papers0 benchmarksImages

Well-being Dataset (Cambridge Well-being Dataset for Psychological Distress Analysis)

The dataset is a private dataset collected for automatic analysis of psychological distress. It contains self-reported distress labels provided by human volunteers. The dataset consists of 30-min interview recordings of participants.

1 papers1 benchmarksAudio, Speech, Time series, Videos

XA Bin-Picking

XA Bin-Picking is a point-cloud dataset comprising both simulated and real-world scenes with three industrial parts. Synthesized scenes consists of 1000 training samples. The test samples are real scenes and the ground truth instance labels are made manually. There are 20 to 30 identical types of parts randomly piled up in a scene. Each scene contains about 60,000 boundary points. Each point in the scene has instance annotations. The parts are texture-less and have no discernible color. Both of training samples and test sam- ples only contain the boundary points of parts.

1 papers0 benchmarks3D

Monkey V1 dataset

This dataset is used for neural co-training. mtl_monkey_dataset: was used for our MTL-Monkey model and involves neural responses that were predicted by a single-task trained model on real monkey V1. mtl_oracle_dataset: was used for our MTL-Oracle model and involves neural responses that were predicted by our image classification oracle. mtl_shuffled_dataset: was used for our MTL-Shuffled model and is the result of shuffling the mtl_monkey_dataset across images.

1 papers0 benchmarks

BIDCD (Bosch Industrial Depth Completion Dataset)

Bosch Industrial Depth Completion Dataset (BIDCD) is an RGBD dataset for of static table-top scenes with industrial objects. The data was collected with a RealSense depth-camera mounted on a robotic arm, i.e. from multiple Points-of-View (POV), approximately 60 for each scene. We generated depth ground truth with a customized pipeline for removing erroneous depth values, and applied Multi-View geometry to fuse the cleaned depth frames and fill-in missing information. The fused scene mesh was back-projected to each POV, and finally a bi-lateral filter was applied to reduce the remaining holes.

1 papers0 benchmarksImages, RGB-D

Bambara Language Dataset

A Bambara dialectal dataset dedicated for Sentiment Analysis, available freely for Natural Language Processing research purposes

1 papers0 benchmarksTexts

TrUMAn (Trope Understanding in Movies and Animations)

Trope Understanding in Movies and Animations (TrUMAn) is a dataset intending to evaluate and develop learning systems beyond visual signals.

1 papers0 benchmarksVideos

OpenStreetMap Multi-Sensor Scene Classification

A high-resolution multi-sensor remote sensing scene classification dataset, appropriate for training and evaluating image classification models in the remote sensing domain.

1 papers0 benchmarksHyperspectral images, Images
PreviousPage 402 of 1000Next