19,997 machine learning datasets
19,997 dataset results
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
The M3AV (Multimodal, Multigenre, and Multipurpose Audio-Visual) is a novel dataset proposed for academic lectures¹. It contains almost 367 hours of videos from five sources covering topics in computer science, mathematics, and medical and biology¹⁴.
COCO-N Medium introduces a stochastic benchmark that simulates common real-world scenarios with noticeable label inaccuracies in the COCO dataset. This benchmark combines class and spatial noises to create a challenging yet realistic evaluation framework for instance segmentation models. It mimics datasets manually annotated by crowd workers, where a moderate level of label noise is expected. By incorporating both class and spatial inaccuracies, COCO-N Medium allows researchers to assess their models' basic robustness to label noise, providing insights into performance in typical real-world applications where perfect annotations are rare. This medium-level benchmark serves as a crucial middle ground, offering a more rigorous test than minimally noisy datasets while remaining within the bounds of commonly encountered data quality issues. COCO-N Medium enables a nuanced evaluation of model performance under realistic conditions, helping identify areas for improvement in handling noisy la
The MAPLE benchmark constructed by us contains 20 datasets across 19 fields for scientific literature tagging. It also has a graph format, which can be used for graph mining tasks (e.g., node classification, link prediction). Refer to its homepage for more details.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Tecnalia Hyperspectral Dataset contains different non-ferreous fractions of Waste from Electric and Electronic Equipment (WEEE) of Copper, Brass, Aluminum, Stainless Steel and White Copper. Images were captured by a hyperspectral Specim PHF Fast10 camera that is able to capture wavelengths in the range 400 to 1000 nm with a spectral resolution of less than 1 nm. The PHF Fast10 camera is equipped with a CMOS sensor (1024 × 1024 resolution), a Camera Link interface and a special Fore objective OL10. The provided dataset contains 76 uniformly distributed wave-lengths in the spectral range [415.05 nm, 1008.10 nm]. Illumination setup, as described in \cite{picon2012real}, was specifically designed to reduce the specular reflections generated by the surface of the non-ferrous materials and to provide a homogeneous and even illumination that covers the wavelengths sensitive to the hyperspectral camera. The illumination system consists of a parabolic surface that uniformly distributes the lig
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Learning dynamical systems that can generalize to various parameter changes in their underlying ODEs or PDEs is a significant but challenging task. In recent years, several methods have been proposed to learn such systems. This repository contains a collection of datasets that can be used to benchmark these methods. We aim to provide a consistent interface to access each dataset.
VSTaR-1M is a 1M instruction tuning dataset, created using Video-STaR, with the source datasets: * Kinetics700 * STAR-benchmark * FineDiving
Music recommendation for videos attracts growing interest in multi-modal research. However, existing systems focus primarily on content compatibility, often ignoring the users’ preferences. Their inability to interact with users for further refinements or to provide explanations leads to a less satisfying experience. We address these issues with MuseChat, a first-of-its-kind dialogue-based recommendation system that personalizes music suggestions for videos. Our system consists of two key functionalities with associated modules: recommendation and reasoning. The recommendation module takes a video along with optional information including previous suggested music and user’s preference as inputs and retrieves an appropriate music matching the context. The reasoning module, equipped with the power of Large Language Model (Vicuna-7B) and extended to multi-modal inputs, is able to provide reasonable explanation for the recommended music. To evaluate the effectiveness of MuseChat, we build
This repository includes the experimental dataset acquired to evaluate our radar-leg odometry algorithm, Co-RaL, accepted by IEEE IROS 2024. Our dataset includes sensor data of chip radar, imu, velodyne, and kinematic data from Boston Dynamics SPOT (Joint encoders and contact sensors). Each sequence is acquired with different environments to evaluate the algorithm performance generally. The dataset is provided with ROS Bag file format.
The PQAref dataset is a dataset for fine-tuning large language models for referenced question-answering in biomedical domain.
The FrodoBots 2K Dataset is a diverse collection of camera footage, GPS, IMU, audio recordings & human control data collected from ~2,000 hours of tele-operated sidewalk robots driving in 10+ cities.
A new SIS benchmark designed to assess generation performance under noisy conditions, simulating human error that can occur during real-world applications. [DS] employs downsampled semantic maps that are resized by nearest-neighbor interpolation, simulating human errors e.g., jagged edges and coarse/low resolution user inputs.
A new SIS benchmark designed to assess generation performance under noisy conditions, simulating human error that can occur during real-world applications. [Edge] masks the edges of instances with an unlabeled class, imitating incomplete annotations around edges, especially between instances.
A new SIS benchmark designed to assess generation performance under noisy conditions, simulating human error that can occur during real-world applications. [Random] randomly adds an unlabeled class to the semantic maps, mimicing unintended user error and extreme random noise.
The dataset contains 3 million attribute-value annotations across 1257 unique categories created from 2.2 million cleaned Amazon product profiles. It is a large, multi-sourced, diverse dataset for product attribute extraction study.
MIR-ST500 Good for the following task: Singing transcription (singing pitch to music note conversion) Used in several papers published by Roger Jang's Lab
We redistribute a suite of datasets as part of the YourMT3 project. The license for redistribution is attached.
APPROVE consists of curated YouTube videos annotated with educational content. APPROVE consists of 193 hours of expert-annotated videos with 19 classes (7 literacy codes, 11 math, and background) and each video is associated with approximately 3 labels on average.