19,997 machine learning datasets
19,997 dataset results
ShuttleSet22 is a badminton singles dataset which is collected from high-ranking matches in 2022. ShuttleSet22 consists of 30,172 strokes in 2,888 rallies in the training set, 1,400 strokes in 450 rallies in validation set, and 2,040 strokes in 654 rallies in testing set with detailed stroke-level metadata within a rally.
The RoseBlooming dataset is a stage-specific flower dataset for detection. The dataset, consisting of overhead images, contains two rose cultivars and was filmed over a period of months. The dataset has 519 images, and most of the images contain several bounding boxes. Therefore, this dataset contains over 7,000 bounding boxes. The developmental stages of flowering branches were visually classified and annotated into two stages: rose_small, and rose_large. For the rose variation, the dataset includes 2 rose cultivars (‘Samourai 08’ and ‘Blossom Pink’ roses). The dataset contains images under various weather conditions.
Dataset Introduction. We create this dataset to test the working memory capacity of language models. We choose the N-back task because it is widely used in cognitive science as a measure of working memory capacity. To create the N-back task dataset, we generated 30 blocks of trials for $N = {1, 2, 3}$, respectively. Each block contains 30 trials, including 10 match trials and 20 nonmatch trials. The dataset for each block is stored in a text file. The first line in the text file is the letter presented on every trial. The second line is the condition corresponding to every letter in the first line ('m':this is a match trial; '-': this is a nonmatch trial). We have created many versions of the N-back task, including verbal ones and spatial ones.
Filip Milojkovic, August 13, 2021, "GEM HOUSE openData: German Electricity consumption in Many HOUSEholds over three years 2018-2020 (Fresh Energy)", IEEE Dataport, doi: https://dx.doi.org/10.21227/4821-vf03.
The UTRSet-Real dataset is a comprehensive, manually annotated dataset specifically curated for Printed Urdu OCR research. It contains over 11,000 printed text line images, each of which has been meticulously annotated. One of the standout features of this dataset is its remarkable diversity, which includes variations in fonts, text sizes, colours, orientations, lighting conditions, noises, styles, and backgrounds. This diversity closely mirrors real-world scenarios, making the dataset highly suitable for training and evaluating models that aim to excel in real-world Urdu text recognition tasks.
The UTRSet-Synth dataset is introduced as a complementary training resource to the UTRSet-Real Dataset, specifically designed to enhance the effectiveness of Urdu OCR models. It is a high-quality synthetic dataset comprising 20,000 lines that closely resemble real-world representations of Urdu text.
The UrduDoc Dataset is a benchmark dataset for Urdu text line detection in scanned documents. It is created as a byproduct of the UTRSet-Real dataset generation process. Comprising 478 diverse images collected from various sources such as books, documents, manuscripts, and newspapers, it offers a valuable resource for research in Urdu document analysis. It includes 358 pages for training and 120 pages for validation, featuring a wide range of styles, scales, and lighting conditions. It serves as a benchmark for evaluating printed Urdu text detection models, and the benchmark results of state-of-the-art models are provided. The Contour-Net model demonstrates the best performance in terms of h-mean.
FLIP includes several benchmark datasets that contain a variety of protein sequences, each with a real-valued label indicating its "fitness" (how well the protein performs some particular function). The goal is to predict the fitness of a given protein sequence using the sequence. Different representations of protein sequences (e.g. learned embeddings from large language models) may prove helpful here.
Introduction NBMOD is a dataset created for researching the task of specific object grasp detection by robots in noisy environments. The dataset comprises three subsets: Simple background Single-object Subset (SSS), Noisy background Single-object Subset (NSS), and Multi-Object grasp detection Subset (MOS). The SSS subset contains 13,500 images, the NSS subset contains 13,000 images, and the MOS subset contains 5,000 images.
VFD-2000 is a video fight detection dataset containing more than 2000 videos. YouTube is the data source. Specific scenarios are searched using “fight” as a search keyword, for example, “street fight”, “beach fight”, and “violence in the restaurant”. 200 videos under 20 different scenes are collected.
Despite the successes of recent developments in visual AI, different shortcomings still exist; from missing exact logical reasoning, to abstract generalization abilities, to understanding complex and noisy scenes. Unfortunately, existing benchmarks, were not designed to capture more than a few of these aspects. Whereas deep learning datasets focus on visually complex data but simple visual reasoning tasks, inductive logic datasets involve complex logical learning tasks, however, lack the visual component. To address this, we propose the visual logical learning dataset, V-LoL, that seamlessly combines visual and logical challenges. Notably, we introduce the first instantiation of V-LoL, V-LoL-Train, -- a visual rendition of a classic benchmark in symbolic AI, the Michalski train problem. By incorporating intricate visual scenes and flexible logical reasoning tasks within a versatile framework, V-LoL-Train provides a platform for investigating a wide range of visual logical learning chal
Used in the development of Topographs: Topological Reconstruction of Particle Physics Processes using Graph Neural Networks
DiaSafety is a comprehensive dialogue safety dataset. It consists of 11K contextual dialogues under 7 unsafe subaspects in chitchat.
A dataset made of 3D image data and their embeddings to test TomoSAM
Dissonance Twitter Dataset is a dataset collected from annotating tweets for dissonance.
It consists of 32x32 pixel images of shapes with multiple attributes (size, location, rotation, color). Each image is also paired with its ground truth information (attributes), and a natural language description (English) of the image.
Air pollution management through wind speed forecasting: the time series exhibits a daily cyclical behavior and a long-term seasonality.
This new dataset represents a subset of the ImageNet1k. It consists of 99000 images and 150 classes. 90000 of them are for training, 600 images for each class. The validation test size is 7500. For testing, we add 1500 images from the ImageNetV2 Top-Images dataset to the validation.
The dataset aims to provide system prompts and user prompts for assistant. You should make random pairs and compute human preference for both system prompt obedience and user prompt relevance through A/B testing.
This is the replication data for the paper: "Crossing the Linguistic Causeway: Ethnonational Differences on Soundscape Attributes in Bahasa Melayu".