19,997 machine learning datasets
19,997 dataset results
Dataset of 374 photos of hand-drawn sketches of App Inventor apps used for development of the Sketch2aia model for automatic generation of App Inventor wireframes from hand-drawn sketches.
In NLP, text classification is one of the primary problems we try to solve and its uses in language analyses are indisputable. The lack of labeled training data made it harder to do these tasks in low resource languages like Amharic. The task of collecting, labeling, annotating, and making valuable this kind of data will encourage junior researchers, schools, and machine learning practitioners to implement existing classification models in their language. In this short paper, we aim to introduce the Amharic text classification dataset that consists of more than 50k news articles that were categorized into 6 classes. This dataset is made available with easy baseline performances to encourage studies and better performance experiments.
THEOStereo is a dataset providing synthetic stereo image pairs and their corresponding scene depth and will be published along with 1. All images follow the omnidirectional camera model. In total, there are 31,250 omnidirectional images pairs. The training set contains 25,000 image pairs. For validation and testing there are 3,125 image pairs, respectively. For each pair, there is a ground truth depth map describing the pixel-wise distance of the object along the left camera's z-axis. The virtual omnidirectional cameras exhibit a FOV of 180 degrees and can be described using Kannala's camera model 2. The distortion parameters are k_1 = 1 and k_2 = k_3 = k_4 = k_5 = 0. The length of the stereo camera's baseline was 0.3 AU (approx. 15 cm, not 30 cm!). Please do not forget to cite 1 if you use the dataset in your work. Thank you.
The VIriors Action Recognition Challenge uses a subset of the UCF101 action recognition dataset:
Tsinghua Dogs is a fine-grained classification dataset for dogs, over 65% of whose images are collected from people's real life. Each dog breed in the dataset contains at least 200 images and a maximum of 7,449 images, basically in proportion to their frequency of occurrence in China, so it significantly increases the diversity for each breed over existing dataset. Furthermore, Tsinghua Dogs annotated bounding boxes of the dog’s whole body and head in each image, which can be used for supervising the training of learning algorithms as well as testing them.
The data set consists of 6257 labeled images of Bose-Einstein condensates (BECs) with and without solitonic excitations, including kink solitons and solitonic vortices. Each element of the data set contains a masked image (132x164 pixels) of 2D atomic density used to train the machine learning model used in the paper "Machine-learning enhanced dark soliton detection in Bose-Einstein condensates," (https://arxiv.org/abs/2101.05404), and a label indicating the class a given image belongs to (0 indicates no solitons, 1 indicates a single soliton, and 2 indicates other excitations). The data structure file and project description are included with the data. This data set was used to train a deep convolutional neural network to automatically recognize whether or not a lone dark soliton has been created in BECs that was then implemented within an automated soliton detection and positioning system (see https://arxiv.org/abs/2101.05404 for details).
The ConScenD dataset consists of over 340 scenarios extracted from the naturalistic highway dataset highD. This scenarios can be used to test for the introduction of Level 3 Automated Lane Keeping Systems according to the UNECE R157 ALKS Regulation.
Darija Open Dataset (DODa) is an open-source project for the Moroccan dialect. With more than 10,000 entries DODa is arguably the largest open-source collaborative project for Darija-English translation built for Natural Language Processing purposes. In fact, besides semantic categorization, DODa also adopts a syntactic one, presents words under different spellings, offers verb-to-noun and masculine-to-feminine correspondences, contains the conjugation of hundreds of verbs in different tenses, and many other subsets to help researchers better understand and study Moroccan dialect.
Levantine Twitter dataset for Misogynistic language (LeT-Mi) is an Arabic Levantine Twitter dataset for misogynistic language to be the first benchmark dataset for Arabic misogyny.
HW-NAS-Bench is a dataset for HardWare-aware Neural Architecture Search (HW-NAS). It is the first dataset for HW-NAS research aiming to democratize HW-NAS research to non-hardware experts and facilitate a unified benchmark for HW-NAS to make HW-NAS research more reproducible and accessible, covering two SOTA NAS search spaces including NAS-Bench-201 and FBNet
The dataset is derived from the MSK-IMPACT dataset designed and published by Zehir using the code published by Penson et al.. The derivation process is described in Development of Genome-Derived Tumor Type Prediction to Inform Clinical Cancer Care.
Data from: Using network approaches to enhance the analysis of cross-linguistic polysemies
This is a benchmark for neural paraphrase detection, to differentiate between original and machine-generated content.
Video class agnostic segmentation (VCAS) is the task of segmenting objects without regards to its semantics combining appearance, motion and geometry from monocular video sequences. The main motivation behind this is to account for unknown objects in the scene and to act as a redundant signal along with the segmentation of known classes for better safety as shown in the following Figure.
The Universal-Scale object detection Benchmark (USB) is a benchmark for object detection that has variations in object scales and image domains by incorporating COCO with the recently proposed Waymo Open Dataset and Manga109-s dataset. To enable fair comparison, USB establishes different protocols by defining multiple thresholds for training epochs and evaluation image resolutions.
The dataset contains anonymised hotel checkins. The dataset contains train and test parts, the in the test part city of the last checkin is masked. The goal is to predict this masked checkin.
The UJIIndoorLoc is a Multi-Building Multi-Floor indoor localization database to test Indoor Positioning System that rely on WLAN/WiFi fingerprint.
Cell-200 is a a dataset of synthetic fluorescence microscopy images with cell populations generated by SIMCEP. The Cell-200 dataset consists of 200,000 $64\times 64$ grayscale images. The number of cells per image ranges from 1 to 200 and there are 1,000 images for each cell count. However, only a subset of Cell-200 with only odd cell counts and 10 images per count (1,000 training images in total) is used for the GAN training.