TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Operator learning using random features: a tool for scientific computing

This repository contains the datasets corresponding to the two benchmark problems appearing in the SIAM papers "The Random Feature Model for Input-Output Maps between Banach Spaces" [SIAM J. Sci. Comput., 43 (2021), pp. A3212–A3243] and the paper "Operator learning using random features: a tool for scientific computing" [to appear in SIAM Review (2024)].

1 papers0 benchmarks

An operator learning perspective on parameter-to-observable maps

This repository contains the datasets corresponding to the three benchmark problems for the Fourier Neural Mappings scientific machine learning architectures. The first file is the data for the advection-diffusion problem, the second for the airfoil problem, and the third for the elliptic homogenization materials problem.

1 papers0 benchmarks

MedLFQA (Medical Long-form Question Answering)

MedLFQA is reconstructed by reformulating the current four biomedical long-form question-answering benchmark datasets: LiveQA, MedicationQA, HealthsearchQA, and K-QA. MedLFQA consists of four components: question (Q), answer (A), must-have statements (MH), and nice-to-have statements (NH). It facilitates the automatic evaluation of models' responses and provides a comprehensive understanding of how the model responds to a patient's question.

1 papers0 benchmarks

YADL (Yet Another Data LAke)

Files composing the YADL data lake, for the paper "Retrieve, Merge, Predict: Augmenting Tables with Data Lakes (Experiment, Analysis & Benchmark Paper)"

1 papers0 benchmarksTabular

ECLAIR (ECLAIR: A High-Fidelity Aerial LiDAR Dataset for Semantic Segmentation)

ECLAIR (Extended Classification of Lidar for AI Recognition), a new outdoor large-scale aerial LiDAR dataset designed specifically for advancing research in point cloud semantic segmentation. As the most extensive and diverse collection of its kind to date, the dataset covers a total area of 10km2 with close to 600 million points and features eleven distinct object categories. To guarantee the dataset's quality and utility, we have thoroughly curated the point labels through an internal team of experts, ensuring accuracy and consistency in semantic labeling. The dataset is engineered to move forward the fields of 3D urban modeling, scene understanding, and utility infrastructure management by presenting new challenges and potential applications.

1 papers6 benchmarks3D, Point cloud

AUR & UMB dataset (Anticancer Efficacy of Auraptene & Umbelliprenin: In Vitro Viability Dataset)

This dataset contains quantitative data on the anticancer effects of the natural coumarins Auraptene (AUR) and Umbelliprenin (UMB) across 27 studies. The data were collected from published literature reporting the impacts of AUR and UMB treatment on the viability of diverse human cancer cell lines.

1 papers1 benchmarksBiology

JAZZVAR Dataset (JAZZVAR: A Dataset of Variations found within Solo Piano Performances of Jazz Standards for Music Overpainting)

Jazz pianists often uniquely interpret jazz standards. Passages from these interpretations can be viewed as sections of variation. We manually extracted such variations from solo jazz piano performances. The JAZZVAR dataset is a collection of 502 pairs of Variation and Original MIDI segments. Each Variation in the dataset is accompanied by a corresponding Original segment containing the melody and chords from the original jazz standard. Our approach differs from many existing jazz datasets in the music information retrieval (MIR) community, which often focus on improvisation sections within jazz performances. In this paper, we outline the curation process for obtaining and sorting the repertoire, the pipeline for creating the Original and Variation pairs, and our analysis of the dataset. We also introduce a new generative music task, Music Overpainting, and present a baseline Transformer model trained on the JAZZVAR dataset for this task. Other potential applications of our dataset inc

1 papers0 benchmarks

AlpacaEval-TH

AlpacaEval in Thai.

1 papers0 benchmarksTexts

MT-Bench-TH

MT-Bench in Thai.

1 papers0 benchmarksTexts

MultiSenseBadminton (MultiSenseBadminton: Wearable Sensor–Based Biomechanical Dataset for Evaluation of Badminton Performance)

The sports industry is witnessing an increasing trend of utilizing multiple synchronized sensors for player data collection, enabling personalized training systems with multi-perspective real-time feedback. Badminton could benefit from these various sensors, but there is a scarcity of comprehensive badminton action datasets for analysis and training feedback. Addressing this gap, this paper introduces a multi-sensor badminton dataset for forehand clear and backhand drive strokes, based on interviews with coaches for optimal usability. The dataset covers various skill levels, including beginners, intermediates, and experts, providing resources for understanding biomechanics across skill levels. It encompasses 7,763 badminton swing data from 25 players, featuring sensor data on eye tracking, body tracking, muscle signals, and foot pressure. The dataset also includes video recordings, detailed annotations on stroke type, skill level, sound, ball landing, and hitting location, as well as s

1 papers0 benchmarksRGB Video, Time series

MSNER (Multilingual Spoken Named Entity Recognition)

This dataset contains named entities annotations for European Parliament recordings in Dutch, French, German and Spanish. The entity annotation scheme follows OntoNotes v5. The original unannotated dataset is VoxPopuli.

1 papers0 benchmarksSpeech, Texts

PaRoutes

We introduce a framework for benchmarking multi-step retrosynthesis methods, i.e. route predictions, called PaRoutes. The framework consists of two sets of 10 000 synthetic routes extracted from the patent literature, a list of stock compounds, and a curated set of reactions on which one-step retrosynthesis models can be trained

1 papers0 benchmarksTexts

Reglamento_Aeronautico_Colombiano_2024

Dataset Details Total Labeled: 100%

1 papers0 benchmarksTexts

Cost awareness in commit messages and issue trackers

Dataset of commit messages and issues containing evidence of cost awareness.

1 papers0 benchmarks

Cost awareness in Stack Overflow discussions

Dataset of Stack Overflow questions about Terraform with cost-related keywords.

1 papers0 benchmarks

RTE3-FR

RTE3-FR dataset is the French translation of the Textual Entailment English dataset used in the RTE-3 Challenge (https://nlp.stanford.edu/RTE3-pilot).

1 papers0 benchmarksTexts

GQNLI-FR

GQNLI-FR is a manually translated French version of the GQNLI challenge dataset, originally written in English.

1 papers0 benchmarksTexts

Probeable Problems ICER 2024

Functionally correct (ok) and incorrect (buggy) solutions to five Probleable Problems: http://arxiv.org/abs/2405.15123 The ok solutions correspond to attempts that successfully probed all ambiguities in the given specification; the buggy solutions represent attempts that addressed these ambiguities partially A more nuanced analysis (beyond ok/buggy) of these attempts may reveal greater insights

1 papers0 benchmarks

LMOT (Low-light Multi-object Tracking Dataset)

The Low-light Multi-object Tracking Dataset (LMOT) is a large-scale dataset that focuses on multi-object tracking in dark scenes. It consists of two parts: 1) The low-light and well-lit videos captured by our dual-camera system. 2) The real world low-light videos captured by a simple camera, to evaluate the generalization in real night scenarios. The videos are provided in both RAW format and sRGB format. After careful annotation, we collect 32 video sequences (2.3\times MOT17), over 35K frames (3.1 \times MOT17) and over 815K bounding boxes (2.8 \times MOT17).

1 papers0 benchmarks

MultiBypass140

BernBypass70 is a dataset consisting of 70 surgical videos of LRYGB at Inselspital, Bern University Hospital, Switzerland. The videos were recorded at a resolution of 720 × 576 at 25 fps.

1 papers0 benchmarks
PreviousPage 501 of 1000Next