19,997 machine learning datasets
19,997 dataset results
This dataset contains several instances of the Offline Nanosatellite Task Scheduling (ONTS) problem, based on the parameters of the FloripaSat-1 mission. Each instance (.json file) is paired with a (quasi-)optimal solution vector (_opt.npz file) and the 500 best solutions found (_sols.npz file).
ColorSVG-100K contains:
This repository contains data for a research project involving graph neural networks (GNNs) applied to mechanical metamaterials and their deformations.
The ISP-AD Dataset is a large-scale anomaly detection dataset, representing a real-world industrial use case. It contains 312,674 fault-free and 246,375 defective samples, including 245,664 synthetic defects and 711 real defects collected on the factory floor.
dataset for Generative Photography
In this paper, we propose RFUAV as a new benchmark dataset for radio-frequency based (RF-based) unmanned aerial vehicle (UAV) identification and address the following challenges: Firstly, many existing datasets feature a restricted variety of drone types and insufficient volumes of raw data, which fail to meet the demands of practical applications. Secondly, existing datasets often lack raw data covering a broad range of signal-to-noise ratios (SNR), or do not provide tools for transforming raw data to different SNR levels. This limitation undermines the validity of model training and evaluation. Lastly, many existing datasets do not offer open-access evaluation tools, leading to a lack of unified evaluation standards in current research within this field. RFUAV comprises approximately 1.3 TB of raw frequency data collected from 37 distinct UAVs using the Universal Software Radio Peripheral (USRP) device in real-world environments. Through in-depth analysis of the RF data in RFUAV, we
The causal reasoning dataset is generated using the Causal Reasoning in Closed Daily Activities (COLD) framework that helps evaluate large language models (LLMs) on their causal reasoning abilities within real-world, everyday activities. This dataset provides causal questions that simulate common activities such as shopping, baking a cake, riding a bus, planting a tree, and going on a train ride. With approximately 9 million causal queries, the COLD dataset challenges LLMs to understand and reason about the causal relationships between events that are familiar and grounded in human experience.
MICCAI Challenge 2024
The HDRT dataset is a large-scale dataset designed for infrared-guided high dynamic range (HDR) imaging. It includes aligned infrared (IR), standard dynamic range (SDR), and HDR images to facilitate research in multi-modal fusion, HDR imaging, and related areas.
AerialMPT is a dataset for pedestrian tracking in aerial image sequences and presents real-world challenges for MOT algorithms such as low frame rate, small moving objects, and complex backgrounds. AerialMPT consists of 14 sequences and 307 frames with an average size of 425 × 358 pixels. The images were acquired by DLR's 4K camera system from altitudes ranging from 600 m to 1400 m, resulting in spatial resolutions (GSDs) ranging from 8 cm/pixel to 13 cm/pixel. In a post-processing step, the images were co-registered, geo-referenced, and cropped for each region of interest, resulting in sequences of 2 fps. The images were acquired during different flight campaigns between 2016 and 2017, over different scenes containing pedestrians and with different crowd densities and movement complexities.
This dataset contains Material-Point-Method (MPM) simulations for various materials, including water, sand, plasticine, jelly, and rigid collisions. Each material is represented as point-clouds that evolve over time. The dataset is designed for learning and predicting MPM-based physical simulations. Each material contains 50 trajectories with different initial velocity field.
This dataset consists of computer-generated images for gas leakage segmentation. It features diverse backgrounds, interfering foreground objects, and precise ground truth annotations.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
PlainFact is a high-quality human-annotated dataset with fine-grained explanation (i.e., added information) annotations.
A collection of datasets and benchmarks for large-scale Performance Modeling with LLMs.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Pick-a-Filter is a semi-synthetic dataset constructed from Pick-a-Pic v1 to measure the capability of text-to-image models of adapting to heterogeneous preferences. We assign users from V1 randomly into two groups: those who prefer blue, cooler image tones (G1) and those who prefer red, warmer image tones (G2). After constructing this split, we apply the following logic to construct the dataset:
Dataset used to train ZAugNet, a neural network for Z-slice augmentation, that encompasses a variety of shapes, textures, and microscopy techniques, as described below: Ascidian Embryos: This dataset consists of 3D confocal images of P. mammillata embryos, captured using fluorescence microscopy. The plasma membrane was imaged using a PH::Tomato construct, and images were taken at 20°C with a Leica TCS SP8 inverted microscope, resulting in cubic voxel datasets. Credits: Rémi Dumollard, Alex McDougall. Cell Nuclei: This dataset includes 3D confocal images of colorectal cancer organoids, stained with DAPI. The images were captured using a Nikon Spatial Array Confocal (NSPARC) detector with 40x objective, providing high-resolution data on organoid structures. Credits: Yekaterina A. Miroshnikova. Filaments of Microtubules: This dataset features 3D images of microtubules in Mouse Embryonic Fibroblasts, captured using a Zeiss LSM 900 Airyscan2 with a high-resolution 63x oil objective. The ima
A collection of prior text datasets assembled for hypothesis generation. See more info on HuggingFace: https://huggingface.co/datasets/rmovva/HypotheSAEs