19,997 machine learning datasets
19,997 dataset results
Dataset accompanying paper Klein, N., Siegle, J.H., Teichert, T., Kass, R.E. (2021) "Cross-population coupling of neural activity based on Gaussian process current source densities".
Single-mouse Neuropixels recordings (spikes and LFPs) in NWB format. Dataset used in the paper "Cross-population coupling of neural activity based on Gaussian process current source densities" by Klein, N., Siegle, J.H., Teichert, T., and Kass, R.E. (preprint: https://arxiv.org/abs/2104.10070).
This is a corpus of about 500 computer vision datasets, from which the authors sampled 114 dataset publications across different vision tasks and coded for themes through both structured and qualitative content analysis. This work most closely pairs with the following research question: How do dataset developers in CV and NLP research, describe and motivate the decisions that go into their creation?
Steredo Waterdrop is a real-world dataset for research on stereo waterdrop removal. The dataset contains 837 stereo image pairs captured from 129 indoor and outdoor scenes with various waterdrops, disparities, and illumination conditions. We use the ZED 2 stereo camera for data collection.
gENder-IT is an English-Italian challenge set focusing on the resolution of natural gender phenomena by providing word-level gender tags on the English source side and multiple gender alternative translations, where needed, on the Italian target side.
WikiChurches is a dataset for architectural style classification, consisting of 9,485 images of church buildings. Both images and style labels were sourced from Wikipedia. The dataset can serve as a benchmark for various research fields, as it combines numerous real-world challenges: fine-grained distinctions between classes based on subtle visual features, a comparatively small sample size, a highly imbalanced class distribution, a high variance of viewpoints, and a hierarchical organization of labels, where only some images are labeled at the most precise level.
We propose a new benchmark called Human Video Instance Segmentation (HVIS), which focuses on complex real-world scenarios with sufficient human instance masks and identities. Our dataset contains 805 videos with 1447 detailedly annotated human instances. It also includes various overlapping scenes, which integrates into the most challenging video dataset related to humans.
The NMR-POISE paper can be found at: Anal. Chem. 2021, 93 (31), 10735–10739 (DOI: 10.1021/acs.analchem.1c01767).
This dataset collects 88,077 numerical samples of call options on Shanghai Stock Exchange from 2015-02 to 2020-07. After the pre-processing, 83,427 samples remain in the data set. This data set records only original quotation of call options on Shanghai Stock Exchange, and does not include derivative indicators published by stock brokerage firms.
This is the dataset to "Easing the Conscience with OPC UA: An Internet-Wide Study on Insecure Deployments" [In ACM Internet Measurement Conference (IMC ’20)]. It contains our weekly scanning results between 2020-02-09 and 2020-08-31 complied using our zgrab2 extensions, i.e, it contains an Internet-wide view on OPC UA deployments and their security configurations. To compile the dataset, we anonymized the output of zgrab2, i.e., we removed host and network identifiers from that dataset. More precisely, we mapped all IP addresses, fully qualified hostnames, and autonomous system IDs to numbers as well as removed certificates containing any identifiers. See the README file for more information. Using this dataset we showed that 93% of Internet-facing OPC UA deployments have problematic security configurations, e.g., missing access control (on 24% of hosts), disabled security functionality (24%), or use of deprecated cryptographic primitives (25%). Furthermore, we discover several hundred
MHMD (Modern Historical Movies Dataset) is a dataset for old image colorization, built from historical movies. It consists of 1,353,166 images and 42 labels of eras, nationalities, and garment types for automatic colorization from 147 historical movies or TV series made in modern time.
STN PLAD is a high-resolution and real-world image dataset of multiple high-voltage power line components. It has 2,409 annotated objects divided into five classes: transmission tower, insulator, spacer, tower plate, and Stockbridge damper, which vary in size (resolution), orientation, illumination, angulation, and background.
This dataset consisting 500 set of caption, table and coresponding paper page, processed from DocBank.
This repository holds two datasets: one with both the original binaries and the code sections extracted from them (“full dataset”), and one with only the code sections (“only code sections”). The code sections were extracted by carving out sections of the binary that were marked as executable. The binaries were scraped from Debian repositories.
P. vivax (malaria) infected human blood smears with bounding box annotations. The data consists of two classes of uninfected cells (RBCs and leukocytes) and four classes of infected cells (gametocytes, rings, trophozoites, and schizonts).
<Task description: joint learning of coreference resolution and query rewrite>
VerbCL is a dataset that consists of the citation graph of court opinions, which cite previously published court opinions in support of their arguments. In particular, it focuses on the verbatim quotes, i.e., where the text of the original opinion is directly reused.
This is two-hop relation extraction dataset derived from WikiHop dataset [1].
Invisible Mobile Keyboard Dataset contains user initial, age, type of mobile devices, size of the screen, time taken for typing each phrase, and annotation of typed phrases with coordinate values of the typed position (x and y points). The collected dataset is the first and only dataset for a novel IMK decoding task.
This dataset contains images taken from camera traps set up in the Jura and Ain counties in France. We use this dataset to illustrate the training of a deep learning algorithm with application to animal specie sidentification. See more here https://github.com/oliviergimenez/computo-deeplearning-occupany-lynx.