TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Two-probe macaque monkey auditory LFP

Dataset accompanying paper Klein, N., Siegle, J.H., Teichert, T., Kass, R.E. (2021) "Cross-population coupling of neural activity based on Gaussian process current source densities".

1 papers0 benchmarks

Neuropixels single-mouse LFP data

Single-mouse Neuropixels recordings (spikes and LFPs) in NWB format. Dataset used in the paper "Cross-population coupling of neural activity based on Gaussian process current source densities" by Klein, N., Siegle, J.H., Teichert, T., and Kass, R.E. (preprint: https://arxiv.org/abs/2104.10070).

1 papers0 benchmarks

Computer Vision Values Dataset

This is a corpus of about 500 computer vision datasets, from which the authors sampled 114 dataset publications across different vision tasks and coded for themes through both structured and qualitative content analysis. This work most closely pairs with the following research question: How do dataset developers in CV and NLP research, describe and motivate the decisions that go into their creation?

1 papers0 benchmarks

Stereo Waterdrop

Steredo Waterdrop is a real-world dataset for research on stereo waterdrop removal. The dataset contains 837 stereo image pairs captured from 129 indoor and outdoor scenes with various waterdrops, disparities, and illumination conditions. We use the ZED 2 stereo camera for data collection.

1 papers0 benchmarksImages

gENder-IT

gENder-IT is an English-Italian challenge set focusing on the resolution of natural gender phenomena by providing word-level gender tags on the English source side and multiple gender alternative translations, where needed, on the Italian target side.

1 papers0 benchmarksTexts

WikiChurches

WikiChurches is a dataset for architectural style classification, consisting of 9,485 images of church buildings. Both images and style labels were sourced from Wikipedia. The dataset can serve as a benchmark for various research fields, as it combines numerous real-world challenges: fine-grained distinctions between classes based on subtle visual features, a comparatively small sample size, a highly imbalanced class distribution, a high variance of viewpoints, and a hierarchical organization of labels, where only some images are labeled at the most precise level.

1 papers0 benchmarksImages

HVIS Dataset (Human Video Instance Segmentation Dataset)

We propose a new benchmark called Human Video Instance Segmentation (HVIS), which focuses on complex real-world scenarios with sufficient human instance masks and identities. Our dataset contains 805 videos with 1447 detailedly annotated human instances. It also includes various overlapping scenes, which integrates into the most challenging video dataset related to humans.

1 papers0 benchmarks

Raw data for NMR-POISE

The NMR-POISE paper can be found at: Anal. Chem. 2021, 93 (31), 10735–10739 (DOI: 10.1021/acs.analchem.1c01767).

1 papers0 benchmarks

CallOptionBSM

This dataset collects 88,077 numerical samples of call options on Shanghai Stock Exchange from 2015-02 to 2020-07. After the pre-processing, 83,427 samples remain in the data set. This data set records only original quotation of call options on Shanghai Stock Exchange, and does not include derivative indicators published by stock brokerage firms.

1 papers0 benchmarks

Dataset to "Easing the Conscience with OPC UA: An Internet-Wide Study on Insecure Deployments"

This is the dataset to "Easing the Conscience with OPC UA: An Internet-Wide Study on Insecure Deployments" [In ACM Internet Measurement Conference (IMC ’20)]. It contains our weekly scanning results between 2020-02-09 and 2020-08-31 complied using our zgrab2 extensions, i.e, it contains an Internet-wide view on OPC UA deployments and their security configurations. To compile the dataset, we anonymized the output of zgrab2, i.e., we removed host and network identifiers from that dataset. More precisely, we mapped all IP addresses, fully qualified hostnames, and autonomous system IDs to numbers as well as removed certificates containing any identifiers. See the README file for more information. Using this dataset we showed that 93% of Internet-facing OPC UA deployments have problematic security configurations, e.g., missing access control (on 24% of hosts), disabled security functionality (24%), or use of deprecated cryptographic primitives (25%). Furthermore, we discover several hundred

1 papers0 benchmarks

MHMD (Modern Historical Movies Dataset)

MHMD (Modern Historical Movies Dataset) is a dataset for old image colorization, built from historical movies. It consists of 1,353,166 images and 42 labels of eras, nationalities, and garment types for automatic colorization from 147 historical movies or TV series made in modern time.

1 papers0 benchmarksImages

STN PLAD (STN Power Line Assets Dataset)

STN PLAD is a high-resolution and real-world image dataset of multiple high-voltage power line components. It has 2,409 annotated objects divided into five classes: transmission tower, insulator, spacer, tower plate, and Stockbridge damper, which vary in size (resolution), orientation, illumination, angulation, and background.

1 papers5 benchmarksImages

DocBank-TB (DocBank-Table)

This dataset consisting 500 set of caption, table and coresponding paper page, processed from DocBank.

1 papers0 benchmarksTabular, Texts

ISAdetect dataset (ISAdetect binary file and object code dataset)

This repository holds two datasets: one with both the original binaries and the code sections extracted from them (“full dataset”), and one with only the code sections (“only code sections”). The code sections were extracted by carving out sections of the binary that were marked as executable. The binaries were scraped from Debian repositories.

1 papers0 benchmarks

BBBC041 (P. vivax (malaria) infected human blood smears)

P. vivax (malaria) infected human blood smears with bounding box annotations. The data consists of two classes of uninfected cells (RBCs and leukocytes) and four classes of infected cells (gametocytes, rings, trophozoites, and schizonts).

1 papers0 benchmarksImages

MuDoCo_QueryRewrite (The MuDoCo dataset with Query Rewrite Annotations)

<Task description: joint learning of coreference resolution and query rewrite>

1 papers0 benchmarksTexts

VerbCL

VerbCL is a dataset that consists of the citation graph of court opinions, which cite previously published court opinions in support of their arguments. In particular, it focuses on the verbatim quotes, i.e., where the text of the original opinion is directly reused.

1 papers0 benchmarksTexts

THRED (Two-Hop Relation Extraction Dataset)

This is two-hop relation extraction dataset derived from WikiHop dataset [1].

1 papers0 benchmarksTexts

Invisible Mobile Keyboard Dataset

Invisible Mobile Keyboard Dataset contains user initial, age, type of mobile devices, size of the screen, time taken for typing each phrase, and annotation of typed phrases with coordinate values of the typed position (x and y points). The collected dataset is the first and only dataset for a novel IMK decoding task.

1 papers0 benchmarks

Images from camera traps in the Jura and Ain counties (France)

This dataset contains images taken from camera traps set up in the Jura and Ain counties in France. We use this dataset to illustrate the training of a deep learning algorithm with application to animal specie sidentification. See more here https://github.com/oliviergimenez/computo-deeplearning-occupany-lynx.

1 papers0 benchmarks
PreviousPage 403 of 1000Next