TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

CPM-Real

CPM-Real is a dataset consisting of 3895 images representing real - makeup styles.

1 papers0 benchmarksImages

COPA-HR

The COPA-HR dataset (Choice of plausible alternatives in Croatian) is a translation of the English COPA dataset by following the XCOPA dataset translation methodology. The dataset consists of 1000 premises (My body cast a shadow over the grass), each given a question (What is the cause?), and two choices (The sun was rising; The grass was cut), with a label encoding which of the choices is more plausible given the annotator or translator (The sun was rising).

1 papers0 benchmarksTexts

CASP13 MQA

CASP13 MQA is a dataset that contains predicted models for CASP13 targets and their scores.

1 papers0 benchmarks

GE852

GE852 is a dataset of 852 game engine repositories mined from GitHub in two languages, namely Java and C++. The dataset contains metadata of all the mined repositories including commits, pull requests, issues and so on. This dataset can lays the foundation for empirical investigation in the area of game engines.

1 papers0 benchmarks

Cry Wolf

Cry Wolf is a dataset for cyber security analysis tasks. It is an open-access dataset of 73 true and false Intrusion Detection System (IDS) alarms derived from real-world examples of "impossible travel" scenarios.

1 papers0 benchmarks

IITM-Bandersnatch

IITM-Bandersnatch is a dataset to evaluate traffic analysis techniques. The dataset comprises of data points of the form {encrypted traces, ground truth choices}. To collect each data point, we asked the viewer to watch Bandersnatch from the beginning and note down the choices made by them. At the same time, we collected the encrypted network traffic. As of now, our dataset contains information corresponding to 100 viewers who volunteered for this study.

1 papers0 benchmarks

Alexa Domains

This dataset is composed of the URLs of the top 1 million websites. The domains are ranked using the Alexa traffic ranking which is determined using a combination of the browsing behavior of users on the website, the number of unique visitors, and the number of pageviews. In more detail, unique visitors are the number of unique users who visit a website on a given day, and pageviews are the total number of user URL requests for the website. However, multiple requests for the same website on the same day are counted as a single pageview. The website with the highest combination of unique visitors and pageviews is ranked the highest

1 papers0 benchmarks

MAI (Multi-scene Aerial Image)

MAI is a dataset for multi-scene recognition in single aerial images. It consists of 3,923 labelled large-scale images from Google Earth imagery that covers the United States, Germany, and France. The size of each image is 512 ×512, and spatial resolutions vary from 0.3 m/pixel to 0.6 m/pixel. After capturing aerial images, multiple scene-level labels were manually assigned to each image from in total 24 scene categories, including apron, baseball, beach, commercial, farmland, woodland, parking lot, port, residential, river, storage tanks, sea, bridge, lake, park, roundabout, soccer field, stadium, train station, works, golf course, runway, sparse shrub, and tennis court

1 papers0 benchmarksImages

MLDS (Machine Learning Datasets)

MLDS is a collection of thousands of trained neural networks labelled with the data used to train them. MLDS allows meta weight-space analysis across thousands of networks trained with identical or similar training data.

1 papers0 benchmarks

Continuous Defect Prediction

Continuous Defect Prediction (CDP) is a dataset of more than 11 million data rows, representing files involved in Continuous Integration (CI) builds, that synthesize the results of CI builds with data mined from software repositories. The dataset embraces 1,265 software projects, 30,022 distinct commit authors and several software process metrics that in earlier research appeared to be useful in software defect prediction. In this particular dataset the authors used TravisTorrent as the source of CI data. TravisTorrent synthesizes commit level information from the Travis CI server and GitHub open-source projects repositories.

1 papers0 benchmarks

Dataset of Dockerfiles

This dataset of approximately 178,000 unique Dockerfiles collected from GitHub to facilitate sophisticated semantics-aware static analysis of Dockerfiles. To enhance the usability of this data, the authors use five representations for working with, mining from, and analyzing these Dockerfiles. Each Dockerfile representation builds upon the previous ones, and the final representation, created by three levels of nested parsing and abstraction, makes tasks such as mining and static checking tractable.

1 papers0 benchmarks

Credibility Factors 2020

This dataset focuses on 50 articles about climate science, which were annotated completely by 49 students, 26 Upwork workers, 3 science and 3 journalism experts.

1 papers0 benchmarksTexts

Acticipate

Acticipate is a publicly available dataset with recordings of human body-motion and eye-gaze, acquired in an experimental scenario with an actor interacting with three subjects. It contains synchronised and labelled video+gaze and body motion in a dyadic scenario of interaction.

1 papers0 benchmarksVideos

APND (Arm Point Nav Dataset)

APND (Arm Point Nav Dataset) is a dataset for the generalizable object manipulation task called ARMPOINTNAV, which consists on moving an object in the scene from a source location to a target location.

1 papers0 benchmarks

Mid-level perceptual musical features

This dataset contains annotations for 5000 music files on the following music properties:

1 papers0 benchmarksMusic

JoCAD

JoCAD is a dataset for anomaly detection in citation networks.

1 papers0 benchmarksGraphs

Hand Poses

This is a dataset for benchmarking in-hand manipulation on different robot platforms.

1 papers0 benchmarks

CD18 (Cellphone Dataset with 18 Features)

1 papers4 benchmarks

SumeCzech-NER

SumeCzech-NER contains named entity annotations of SumeCzech 1.0, a Czech news-based summarization dataset.

1 papers0 benchmarksTexts

AskUbuntu

AskUbuntu question dataset is a preprocessed collection of questions taken from the AskUbuntu.com 2014 corpus dump. It also comes with 400*20 manual annotations, marking pairs of questions as "similar" or "non-similar".

1 papers0 benchmarks
PreviousPage 388 of 1000Next