19,997 machine learning datasets
19,997 dataset results
This repository contains the datasets corresponding to the two benchmark problems appearing in the SIAM papers "The Random Feature Model for Input-Output Maps between Banach Spaces" [SIAM J. Sci. Comput., 43 (2021), pp. A3212–A3243] and the paper "Operator learning using random features: a tool for scientific computing" [to appear in SIAM Review (2024)].
This repository contains the datasets corresponding to the three benchmark problems for the Fourier Neural Mappings scientific machine learning architectures. The first file is the data for the advection-diffusion problem, the second for the airfoil problem, and the third for the elliptic homogenization materials problem.
MedLFQA is reconstructed by reformulating the current four biomedical long-form question-answering benchmark datasets: LiveQA, MedicationQA, HealthsearchQA, and K-QA. MedLFQA consists of four components: question (Q), answer (A), must-have statements (MH), and nice-to-have statements (NH). It facilitates the automatic evaluation of models' responses and provides a comprehensive understanding of how the model responds to a patient's question.
Files composing the YADL data lake, for the paper "Retrieve, Merge, Predict: Augmenting Tables with Data Lakes (Experiment, Analysis & Benchmark Paper)"
ECLAIR (Extended Classification of Lidar for AI Recognition), a new outdoor large-scale aerial LiDAR dataset designed specifically for advancing research in point cloud semantic segmentation. As the most extensive and diverse collection of its kind to date, the dataset covers a total area of 10km2 with close to 600 million points and features eleven distinct object categories. To guarantee the dataset's quality and utility, we have thoroughly curated the point labels through an internal team of experts, ensuring accuracy and consistency in semantic labeling. The dataset is engineered to move forward the fields of 3D urban modeling, scene understanding, and utility infrastructure management by presenting new challenges and potential applications.
This dataset contains quantitative data on the anticancer effects of the natural coumarins Auraptene (AUR) and Umbelliprenin (UMB) across 27 studies. The data were collected from published literature reporting the impacts of AUR and UMB treatment on the viability of diverse human cancer cell lines.
Jazz pianists often uniquely interpret jazz standards. Passages from these interpretations can be viewed as sections of variation. We manually extracted such variations from solo jazz piano performances. The JAZZVAR dataset is a collection of 502 pairs of Variation and Original MIDI segments. Each Variation in the dataset is accompanied by a corresponding Original segment containing the melody and chords from the original jazz standard. Our approach differs from many existing jazz datasets in the music information retrieval (MIR) community, which often focus on improvisation sections within jazz performances. In this paper, we outline the curation process for obtaining and sorting the repertoire, the pipeline for creating the Original and Variation pairs, and our analysis of the dataset. We also introduce a new generative music task, Music Overpainting, and present a baseline Transformer model trained on the JAZZVAR dataset for this task. Other potential applications of our dataset inc
AlpacaEval in Thai.
MT-Bench in Thai.
The sports industry is witnessing an increasing trend of utilizing multiple synchronized sensors for player data collection, enabling personalized training systems with multi-perspective real-time feedback. Badminton could benefit from these various sensors, but there is a scarcity of comprehensive badminton action datasets for analysis and training feedback. Addressing this gap, this paper introduces a multi-sensor badminton dataset for forehand clear and backhand drive strokes, based on interviews with coaches for optimal usability. The dataset covers various skill levels, including beginners, intermediates, and experts, providing resources for understanding biomechanics across skill levels. It encompasses 7,763 badminton swing data from 25 players, featuring sensor data on eye tracking, body tracking, muscle signals, and foot pressure. The dataset also includes video recordings, detailed annotations on stroke type, skill level, sound, ball landing, and hitting location, as well as s
This dataset contains named entities annotations for European Parliament recordings in Dutch, French, German and Spanish. The entity annotation scheme follows OntoNotes v5. The original unannotated dataset is VoxPopuli.
We introduce a framework for benchmarking multi-step retrosynthesis methods, i.e. route predictions, called PaRoutes. The framework consists of two sets of 10 000 synthetic routes extracted from the patent literature, a list of stock compounds, and a curated set of reactions on which one-step retrosynthesis models can be trained
Dataset Details Total Labeled: 100%
Dataset of commit messages and issues containing evidence of cost awareness.
Dataset of Stack Overflow questions about Terraform with cost-related keywords.
RTE3-FR dataset is the French translation of the Textual Entailment English dataset used in the RTE-3 Challenge (https://nlp.stanford.edu/RTE3-pilot).
GQNLI-FR is a manually translated French version of the GQNLI challenge dataset, originally written in English.
Functionally correct (ok) and incorrect (buggy) solutions to five Probleable Problems: http://arxiv.org/abs/2405.15123 The ok solutions correspond to attempts that successfully probed all ambiguities in the given specification; the buggy solutions represent attempts that addressed these ambiguities partially A more nuanced analysis (beyond ok/buggy) of these attempts may reveal greater insights
The Low-light Multi-object Tracking Dataset (LMOT) is a large-scale dataset that focuses on multi-object tracking in dark scenes. It consists of two parts: 1) The low-light and well-lit videos captured by our dual-camera system. 2) The real world low-light videos captured by a simple camera, to evaluate the generalization in real night scenarios. The videos are provided in both RAW format and sRGB format. After careful annotation, we collect 32 video sequences (2.3\times MOT17), over 35K frames (3.1 \times MOT17) and over 815K bounding boxes (2.8 \times MOT17).
BernBypass70 is a dataset consisting of 70 surgical videos of LRYGB at Inselspital, Bern University Hospital, Switzerland. The videos were recorded at a resolution of 720 × 576 at 25 fps.