19,997 machine learning datasets
19,997 dataset results
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Current autonomous driving algorithms heavily rely on the visible spectrum, which is prone to performance degradation in adverse conditions like fog, rain, snow, glare, and high contrast. Although other spectral bands like near-infrared (NIR) and long-wave infrared (LWIR) can enhance vision perception in such situations, they have limitations and lack large-scale datasets and benchmarks. Short-wave infrared (SWIR) imaging offers several advantages over NIR and LWIR. However, no publicly available large-scale datasets currently incorporate SWIR data for autonomous driving. To address this gap, we introduce the RGB and SWIR Multispectral Driving (RASMD) dataset, which comprises 100,000 synchronized and spatially aligned RGB-SWIR image pairs collected across diverse locations, lighting, and weather conditions. In addition, we provide a subset for RGB-SWIR translation and object detection annotations for a subset of challenging traffic scenarios to demonstrate the utility of SWIR imaging t
GenoTEX (Genomics Data Automatic Exploration Benchmark) is a benchmark dataset for the automated analysis of gene expression data to identify disease-associated genes while considering the influence of other biological factors. It provides analysis code and results for solving a wide range of gene-trait association (GTA) analysis problems, encompassing dataset selection, preprocessing, and statistical analysis, in a pipeline that follows computational genomics standards. The benchmark includes expert-curated annotations from bioinformaticians to ensure accuracy and reliability.
This dataset is a patched version of The Taste & Affect Music Database by D. Guedes et al. It is a set of captions that describe 100 musical pieces and associate with them gustatory keywords on the basis of Guedes findings.
GraspClutter6D is a large-scale real-world dataset for robust object perception and robotic grasping in cluttered environments. It features 1,000 highly cluttered scenes with dense arrangements (average 14.1 objects/scene with 62.6% occlusion), 200 household, industrial, and warehouse objects captured in 75 diverse environment configurations (bins, shelves, and tables), multi-view data from 4 RGB-D cameras (RealSense D415, D435, Azure Kinect, and Zivid One+), and comprehensive annotations including 736K 6D object poses and 9.3 billion feasible robotic grasps for 52K RGB-D images. The dataset provides a challenging testbed for segmentation, 6D pose estimation, and grasp detection algorithms in realistic cluttered scenarios.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Dataset contains light curves of 6 rocket body types from Mini Mega Tortora database (MMT)[^1]. The dataset was created to be used as a benchmark for rocket body light curve classification. For more informations follow the original paper: RoBo6: Standardized MMT Light Curve Dataset for Rocket Body Classification[^2]
dataset in WWW 2019 "DPLink: User Identity Linkage via Deep Neural Network From Heterogeneous Mobility Data". This data is intended for academic use only. Redistribution of this data is not permitted without our explicit permission.
MATLAB code to reproduce results presented in the paper "Topology-Based Reconstruction Prevention for Decentralised Learning".
ML-ready Global Dataset of elevation map. Adapting Copernicus DEM GLO-30 to the Major TOM framework.
The dataset has thousands of time series. The provided code select 40 times series of four different profiles to compare classical models (ARIMA) and deep learning models (LSTM)
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
VOCEdits: A benchmark for precise geometric object-level editing Sample format: (input image, edit prompt, input mask, ground-truth output mask, ...)
GitBugs is a comprehensive and up-to-date dataset comprising over 150,000 bug reports from nine actively maintained open-source projects, including Firefox, Cassandra, and VS Code. GitBugs aggregates data from Github, Bugzilla and Jira issue trackers, offering standardized categorical fields for classification tasks and predefined train/test splits for duplicate bug detection. In addition, it includes exploratory analysis notebooks and detailed project-level statistics, such as duplicate rates and resolution times. GitBugs supports various software engineering research tasks, including duplicate detection, retrieval augmented generation, resolution prediction, automated triaging, and temporal analysis. The openly licensed dataset provides a valuable cross-project resource for bench- marking and advancing automated bug report analysis. Access the data and code at this https URL.
DivShift North American West Coast DivShift Paper | Extended Version | Code
TGB is a collection of challenging and diverse benchmark datasets for realistic, reproducible, and robust machine learning evaluation on temporal graphs. It includes dynamic link and node property prediction tasks and an automated pipeline from dataset downloading, data loading, evaluation, and submission to the TGB leaderboard. TGB 2.0 includes novel datasets for temporal knowledge graphs and temporal heterogeneous graphs.
Project: Discrete-Time Modeling of Interturn Short Circuits in Interior PMSMs
CAShift is the first multiple normality shift-aware Log-Based Anomaly Detection (LAD) dataset specifically designed for cloud systems, which considers different software roles in cloud systems and attack behavior among cloud components.
a new self-annotated CC-ReID dataset named Cloth-Changing Unreal Person.
BlenderGym is the first comprehensive VLM system benchmark for 3D graphics editing. It evaluates VLM systems through code-based 3D reconstruction tasks. BlenderGym consists of 245 hand-crafted Blender scenes across 5 key graphics editing tasks: procedural geometry editing, lighting adjustments, procedural material design, blend shape manipulation, and object placement.