19,997 machine learning datasets
19,997 dataset results
Dataset of sperm head images with expert-classification labels. The dataset contains 1854 sperm head images obtained from six semen smears and classified by three Chilean referent domain experts according to World Health Organization (WHO) criteria, in one of the following classes: normal, tapered, pyriform, small and amorphous. This gold-standard is aimed for use in evaluating and comparing not only known techniques, but also future improvements to present approaches for classification of human sperm heads for semen analysis.
This dataset is a collection of undirected and unweighted LFR benchmark graphs as proposed by Lancichinetti et al. [1]. We generated the graphs using the code provided by Santo Fortunato on his personal website [2], embedded in our evaluation framework [3], with two different parameter sets. Let N denote the number of vertices in the network, then
Context of the data sets The Zooniverse platform (www.zooniverse.org) has successfully built a large community of volunteers contributing to citizen science projects. Galaxy Zoo and the Milky Way Project were hosted there.
From dataset repository for "2020 International BCI Competition": https://osf.io/pq7vb/?view_only=08e7108d89fd42bab2adbd6b98fb683d
Provide:
Extension of the official KITTI'15 dataset. independently moving instance segmentation ground truth to cover all moving objects, not just a selection of cars and vans.
This dataset contains the full set of experimental waveforms that were used to produce the article "Non-Linear Phase Noise Mitigation over Systems using Constellation Shaping", published in the Journal of Lightwave Technology with DOI: 10.1109/JLT.2019.2917308.
Typography-MNIST is a dataset comprising of 565,292 MNIST-style grayscale images representing 1,812 unique glyphs in varied styles of 1,355 Google-fonts. The glyph-list contains common characters from over 150 of the modern and historical language scripts with symbol sets, and each font-style represents varying subsets of the total unique glyphs. The dataset has been developed as part of the Cognitive Type project which aims to develop eye-tracking tools for real-time mapping of type to cognition and to create computational tools that allow for the easy design of typefaces with cognitive properties such as readability.
A Natural Language Resource for Learning to Recognize Misinformation about the COVID-19 and HPV Vaccines.
EUCA dataset description Associated Paper: EUCA: the End-User-Centered Explainable AI Framework
The malnutrition data, from the United Nations Children's Fund data warehouse, include two variables, stunted growth and the prevalence of low birth weight, collected in 77 countries from 1985 to 2019. Stunted growth is defined as the proportion of newborns aging from 0 to 59 months with a low height-for-age measurement (below two standard deviations). The stunted growth data represent a point sparseness case with 4-23 recordings per nation. The low birth weight data are a partial sparseness case, with recordings during 2000-2015 only.
This is the small version of the MuMiN dataset.
This is the medium version of the MuMiN dataset.
This is the large version of the MuMiN dataset.
iFLYTEK and ChangGuang Satellite jointly held the challenge of extracting cultivated land from high-resolution remote sensing images.
A dataset of music videos with continuous valence/arousal ratings as well as emotion tags.
Stack of 2D gray images of glass fiber-reinforced polyamide 66 (GF-PA66) 3D X-ray Computed Tomography (XCT) specimen.
Icon645 is a large-scale dataset of icon images that cover a wide range of objects:
Synthetic Dataset created in AirSim
TraVLR is a synthetic dataset comprising four visio-linguistic reasoning tasks. Each example encodes the scene bimodally such that either modality can be dropped during training/testing with no loss of relevant information. TraVLR's training and testing distributions are also constrained along task-relevant dimensions, enabling the evaluation of out-of-distribution generalisation.