19,997 machine learning datasets
19,997 dataset results
JustLogic is a natural language deductive reasoning dataset. JustLogic is (i) highly complex, capable of generating a diverse range of linguistic patterns, vocabulary, and argument structures; (ii) prior knowledge independent, eliminating the advantage of models possessing prior knowledge and ensuring that only deductive reasoning is used to answer questions; and (iii) capable of in-depth error analysis on the heterogeneous effects of reasoning depth and argument form on model accuracy.
the FloCo dataset that contains 11,884 flowchart images and their corresponding Python codes.
Contains the datasets and distillation labels which were used in our paper Amin, I., Raja, S., Krishnapriyan, A.S. (2024). Towards Fast, Specialized Machine Learning Force Fields: Distilling Foundation Models via Energy Hessians. Accepted to ICLR 2025. arXiv:2501.09009.
We release the datasets to replicate the results of `Coordinated Reply Attacks in Influence Operations: Characterization and Detection'.
We introduce EMMA (Enhanced MultiModal reAsoning), a benchmark targeting organic multimodal reasoning across mathematics, physics, chemistry, and coding. EMMA tasks demand advanced cross-modal reasoning that cannot be solved by thinking separately in each modality, offering an enhanced test suite for MLLMs' reasoning capabilities.
PoTATO is a dataset designed to enhance the detection of floating plastic waste in aquatic environments by leveraging polarimetric imaging. It comprises 12,380 labeled images of plastic bottles, each enriched with detailed polarimetric information. The dataset aims to improve object detection accuracy under challenging outdoor lighting conditions and water surface reflections. By providing raw image data, PoTATO offers researchers the opportunity to explore novel approaches and advance state-of-the-art object detection algorithms. The dataset and associated code are publicly available
Late third instar wing imaginal discs were cultured in Shields and Sang M3 media (Sigma) supplemented with 2% FBS (Sigma), 1% pen/strep (Gibco), 3ng/ml ecdysone (Sigma) and 2ng/ml insulin (Sigma). Wing discs were cultured in 35mm fluorodishes (WPI) under 12mm filters (Millicell), as described in https://doi.org/10.1038%2Fs41567-019-0618-1
Cleaned and preprocessed version of the Kyokushin Karate Motion Dataset by Szczkesna et al. The original dataset and detailed description can be found at here.
This benchmark library is curated and maintained by the IEEE PES Task Force on Benchmarks for Validation of Emerging Power System Algorithms and is designed to evaluate a well established version of the the AC Optimal Power Flow problem. This introductory video and detailed report present the motivations and goals of this benchmark library. In particular, these cases are designed for benchmarking algorithms that solve the following Non-Convex Nonlinear Program,
Question answering over temporal knowledge graphs (TKGs) is crucial for understanding evolving facts and relationships, yet its development is hindered by limited datasets and difficulties in generating custom QA pairs. We propose a novel categorization framework based on timeline-context relationships, along with \textbf{TimelineKGQA}, a universal temporal QA generator applicable to any TKGs. The code is available at: \url{https://github.com/PascalSun/TimelineKGQA} as an open source Python package.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
ShopTC-100K Dataset The ShopTC-100K dataset is collected using TermMiner, an open-source data collection and topic modeling pipeline introduced in the paper:
From the dataset paper: Brain-computer interfaces (BCIs) can restore communication to people who have lost the ability to move or speak. In this study, we demonstrated an intracortical BCI that decodes attempted speaking movements from neural activity in motor cortex and translates it to text in real-time, using a recurrent neural network decoding approach. With this BCI, our study participant, who can no longer speak intelligibly due to amyotrophic lateral sclerosis, achieved a 9.1% word error rate on a 50-word vocabulary and a 23.8% word error rate on a 125,000-word vocabulary.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
ENSeg Dataset Overview This dataset represents an enhanced subset of the ENS dataset. The ENS dataset comprises image samples extracted from the enteric nervous system (ENS) of male adult Wistar rats (Rattus norvegicus, albius variety), specifically from the jejunum, the second segment of the small intestine.
Speech Recognition Dataset for Oromo Language. š Key features of Sagalee: 100 hours of read speech. 283 gender balanced speakers * Covers different dialects in Oromo language * Open source for research
WEB-IDS23 is a network intrusion detection dataset that includes over 12 million flows, categorizing 20 attack types across FTP, HTTP/S, SMTP, SSH, and network scanning activities. This dataset is documented in the paper "Technical Report: Generating the WEB-IDS23 Dataset," which provides insights into the generation, structure, and key characteristics of the dataset.
The Liver-US dataset is a comprehensive collection of high-quality ultrasound images of the liver, including both normal and abnormal cases. This dataset is designed to facilitate research in medical image classification, with a focus on liver-related conditions. It includes a diverse range of ultrasound images acquired from multiple clinical settings, providing a robust foundation for developing and validating machine learning models in medical image analysis. Detailed Dataset Description
The repository hosts the STURM-Flood dataset, an open-access resource designed for flood extent mapping using Sentinel-1 and Sentinel-2 satellite imagery. The dataset comprises 21,602 Sentinel-1 tiles and 2,675 Sentinel-2 tiles, each of size 128āĆā128 pixels at a resolution of 10 meters, along with corresponding water masks covering 60 flood events globally. This curated dataset is optimized for deep learning applications and provides ground-truth data from the Copernicus Emergency Management Service to facilitate robust model development. We invite researchers and developers to utilize this resource for advancing flood mapping techniques in disaster management. For further details on the methodology, results, and implementation, please refer to our study published in Big Earth Data (2096-4471): https://doi.org/10.1080/20964471.2025.2458714 .