19,997 machine learning datasets
19,997 dataset results
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
The SimBEV dataset is a collection of 320 scenes spread across all 11 CARLA maps and contains data from a variety of sensors, including five camera types (RGB, semantic segmentation, instance segmentation, depth, and optical flow), lidar, semantic lidar, radar, GNSS, and IMU, along with 3D object bounding boxes and accurate bird's-eye view (BEV) ground truth. With each scene lasting 16 seconds at a frame rate of 20 Hz, the SimBEV dataset contains 102,400 annotated frames, over 8 million 3D object bounding boxes, and more than 2.5 billion BEV ground truth labels.
Fruits Dataset for Classification About Dataset
About Dataset The file contains 24K unique figure obtained from various Google resources Meticulously curated figure ensuring diversity and representativeness Provides a solid foundation for developing robust and precise figure allocation algorithms Encourages exploration in the fascinating field of feed figure allocation
About Dataset Context With increasing unrest, security cams should be armed with advance technology.
Platinum Benchmarks are benchmarks that are are carefully curated to minimize label errors and ambiguity, allowing us to measure reliability of models.
CCPT is a dataset containing 12.3K triplets of noun phrases, properties, and property types for conceptual combination.
About Dataset The File contains 3D point cloud data of a Fabricate plant with 10 sequences. Each sequence contains 0-19 days data at every growth stage of the specific sequence.
CompMix-IR Dataset Overview:
We introduced a new dataset of clinical report summaries, annotated with structured information across 15 categories. This dataset was created to address the lack of large-scale resources for clinical IE. It also promotes the development of methods tailored to clinical data, helping to improve healthcare provision. The dataset contains 60, 000 annotated English clinical report summaries, from which we translated over 24, 000 examples into German.
Retrieval-based Clinical Decision Support (ReCDS) can aid clinical workflow by providing relevant literature and similar patients for a given patient. However, the development of ReCDS systems has been severely obstructed by the lack of diverse patient collections and publicly available large-scale patient-level annotation datasets. In this paper, we collect a novel dataset of patient summaries and relations called PMC-Patients to benchmark two ReCDS tasks: Patient-to-Article Retrieval (ReCDS-PAR) and Patient-to-Patient Retrieval (ReCDS-PPR). Specifically, we extract patient summaries from PubMed Central articles using simple heuristics and utilize the PubMed citation graph to define patient-article relevance and patient-patient similarity. PMC-Patients contains 167k patient summaries with 3.1 M patient-article relevance annotations and 293k patient-patient similarity annotations, which is the largest-scale resource for ReCDS and also one of the largest patient collections. Human evalu
RLM25 is an evaluation benchmark containing 619 paired examples of research-level natural language mathematical statements and their corresponding Lean formalizations. The examples are drawn from six real-world formalization projects, and each entry includes essential context along with timestamp information to ensure freshness and avoid data contamination. Its primary purpose is to serve as a challenging testbed for assessing autoformalization systems, helping researchers measure improvements in translating complex, research-level mathematics into formal code.
ProofNetVerif is an evaluation benchmark comprising 3,752 entries, each including an informal mathematical statement, its reference formalization, a predicted formalization, and a binary label indicating semantic equivalence. It is designed to assess autoformalization metrics by providing a challenging testbed for both reference-based and reference-free evaluation approaches.
TUMTraffic-VideoQA is a novel dataset designed to understand spatiotemporal video in complex roadside traffic scenarios. The dataset comprises 1,000 videos, featuring 85,000 multiple-choice QA pairs, 2,300 object captioning, and 5,700 object grounding annotations, encompassing diverse real-world conditions such as adverse weather and traffic anomalies. By incorporating tuple-based spatiotemporal object expressions, TUMTraffic-VideoQA unifies three essential tasks—multiple-choice video question answering, referred object captioning, and spatiotemporal object grounding—within a cohesive evaluation framework.
XMIDI is a comprehensive, large-scale symbolic music dataset that includes accurate emotion and genre labels, consisting of 108,023 MIDI files. The average duration of the music pieces is approximately 176 seconds, yielding a total dataset length of around 5,278 hours.
https://arxiv.org/abs/2502.06858
Dataset Description: NBA Team Statistics, Historical Performance & Betting Odds (2015-2019) Overview This dataset contains team-level box score statistics, historical win percentages, and closing betting odds for NBA games from 2015 to 2019. It supports research in sports analytics, predictive modeling, and betting market efficiency.
The PS-Eval Dataset is a suite of polysemous and monosemous contexts extracted and filtered from the WiC dataset. It aims to evaluate the ability of Sparse Autoencoders (SAEs) to disentangle polysemantic activations into monosemantic features within large language models (LLMs).
SensoDat is a dataset of self-driving car simulation data (30K executed simulations). Concretely, it contains:
Data for "Image-based Backbone Reconstruction for Non-Slender Soft Robots" This dataset provides the data for the forthcoming paper "Image-based Backbone Reconstruction for Non-Slender Soft Robots". The backbone reconstruction method used is based on the method described in Hoffmann et al. [1]. The modifications to this method to support the non-slender soft robot in this dataset are described in the forthcoming paper mentioned above. This dataset holds raw images of pressurized and elongated soft robots and the corresponding reconstructed backbones.