TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

French Open Science Monitor Dataset

This dataset contains the publication data underlying the French Open Science Monitor.

1 papers0 benchmarksTabular

PINO-geo_var-plate_stress

This dataset is well-structured for the physics-informed training of Neural operators for varying domain geometry, which provides the FEM results of solving a 2D plate stress problem in a domain geometry shape of a rectangle with four holes of different locations and sizes. The Github of the paper that first uses this dataset is: https://github.com/WeihengZ/PI-GANO.

1 papers0 benchmarks

PINO-geo_var-darcy-polygon

This dataset is well-structured for the physics-informed training of Neural operators for varying domain geometry, which provides the FEM results of solving a darcy problem in a domain geometry shape of a polygon. The Github of the paper that first uses this dataset is: https://github.com/WeihengZ/PI-GANO.

1 papers0 benchmarks

PC-GITA

PC-GITA is a Spanish speech corpus designed to analyze speech impairments in individuals with Parkinson's Disease (PD).

1 papers0 benchmarksAudio

COph100

we introduce COph100, a novel and challenging dataset known as the Comprehensive Ophthalmology Retinal Image Registration dataset for infants with a wide range of image quality issues constituting the public "RIDIRP" database. COph100 consists of 100 eyes, each with 2 to 9 examination sessions, amounting to a total of 491 image pairs carefully selected from the publicly available dataset. We manually labeled the corresponding ground truth image points and provided automatic vessel segmentation masks for each image.

1 papers0 benchmarksImages

3U-VQA (Usual, Unusual and Unknown object scenarios for LVQA with difficulty scoring dataset)

To tackle the challenge of obtaining out-of-distribution (OOD) data for LVQA models, we introduced a novel dataset named 3U-VQA dataset (Usual, Unusual and Unknown object scenarios for LVQA with difficulty scoring dataset). The dataset comprises question and image sets. Each instance in the questions set is associated with a set of features representing the question-related criteria set. The questions and their ground truth answers are written using placeholders for the objects and their features, which can be specified based on the user needs and requirements. When creating the questions, we avoided deliberately binary (Yes/No) questions to prevent potential bias caused by the question's type in the model response.

1 papers0 benchmarks

IUST_PersonReID

The IUST_PersonReID dataset was developed to address limitations in existing person re-identification datasets by including cultural and environmental contexts unique to Islamic countries, especially Iran and Iraq. Unlike common datasets, which don’t reflect the clothing styles common in these regions—such as hijabs and other coverings—the IUST_PersonReID dataset represents this diversity, helping to reduce demographic bias and improve model accuracy. Collected from various real-world settings under different lighting, camera angles, indoor & outdoor, and weather conditions, this dataset provides extensive, overlapping views across multiple cameras. By capturing these unique conditions, IUST_PersonReID offers a valuable resource for developing re-ID models that perform more reliably across diverse environments and populations.

1 papers4 benchmarksImages

PRMBench_Preview

This is the official dataset for PRMBench. PRMBench is a benchmark dataset for evaluating process-level reward models (PRMs). It consists of 6,216 data instances, each containing a question, a solution process, and a modified process with errors. The dataset is designed to evaluate the ability of PRMs to identify fine-grained error types in the solution process. The dataset is annotated with error types and reasons for the errors, providing a comprehensive evaluation of PRMs.

1 papers0 benchmarksTexts

PubChem: Antagonist of Human D 1 Dopamine Receptor: qHTS

he goal of this project is to use high throughput screening approaches to identify and develop novel, highly selective small molecule allosteric modulators of the D1 DAR for use as in vitro and in vivo pharmacological tools and in proof-of-concept experiments in animal models of neuropsychiatric disease. There are three different types of allosteric modulators that we are seeking, two of which will stimulate or augment receptor signaling, allosteric agonists and potentiators, while the third, allosteric antagonists, will attenuate receptor signaling. The present primary screening assay measures compound antagonism after an EC80 addition of dopamine by tracking calcium flux in a force-coupled, inducible Hek293 Trex D1 cell line.

1 papers0 benchmarks

Data for: Neuromorphic weighted sums with magnetic skyrmions

The following experimental data were obtained on lithography devices made of magnetic multilayer tracks and thin tantalum transverse electrodes by Kerr microscopy and anomalous Hall effect measurements. The results, demonstrating the weighted sum operation using magnetic skyrmions, are published in T. da Câmara Santa Clara Gomes et al., Neuromorphic weighted sums with magnetic skyrmions, Nature Electronics (2024). Please find in the README additional information regarding the data files and the variables.

1 papers0 benchmarks

WayveScenes101

bbb

1 papers0 benchmarksImages

SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers

SPIQA Dataset Card Dataset Details Dataset Name: SPIQA (Scientific Paper Image Question Answering)

1 papers0 benchmarksImages, Texts

Supplementary Information for Machine Learning Applications in Archaeological Practices: A Review

This deposit is supplementary material to "Machine Learning Applications in Archaeological Practices: A Review". It contains six different files:

1 papers0 benchmarks

Drone-Anomaly

Drone-Anomaly rovides 37 training video sequences and 22 testing video sequences from 7 different realistic scenes with various anomalous events. There are 87,488 color video frames (51,635 for training and 35,853 for testing) with the size of 640 × 640 at 30 frames per second.

1 papers0 benchmarks

Guitar-TECHS (Guitar Tones/Techniques, Excerpts & Chords Dataset)

Guitar-TECHS is a comprehensive dataset featuring a variety of guitar techniques, musical excerpts, chords, and scales. These elements are performed by diverse musicians across various recording settings. Guitar-TECHS incorporates recordings from two stereo microphones: an egocentric microphone positioned on the performer’s head and an exocentric microphone placed in front of the performer. It also includes direct input recordings and microphoned amplifier outputs, offering a wide spectrum of audio inputs and recording qualities. All signals and MIDI labels are properly synchronized. Its multi-perspective and multi-modal content makes Guitar-TECHS a valuable resource for advancing data-driven guitar research, and to develop robust guitar listening algorithms.

1 papers0 benchmarksAudio, Midi

Translated SNLI Dataset in Marathi

Translated SNLI Dataset in Marathi A translated version of the SNLI dataset in Marathi, designed for Semantic Textual Similarity (STS) tasks. The translations were generated using the model aryaumesh/english-to-marathi.

1 papers2 benchmarksTexts

GerDaLIR (A German Dataset for Legal Information Retrieval)

GerDaLIR The German Dataset for Legal Information Retrieval (GerDaLIR) is a legal information retrieval dataset comprising a large collection of documents, passages and relevance labels. The large amount of training data we provide enables GerDaLIR to be used as a downstream task for German or multilingual language models. The task provided is a precedent retrieval task based on case documents from the open legal information platform Open Legal Data. Relevance labels are derived from references: If a passage contains a reference to one or more available documents, the passage is used as a query while the referenced cases are labelled as relevant.

1 papers0 benchmarks

GRND (Gramophone Recording Noise Dataset)

Dataset of noise segments extracted from gramophone recordings. It contains 139 min of noises, divided into 2430 segments from 1386 different recordings, dated between 1902 and 1966.

1 papers0 benchmarks

PlaningItByEar

See the description in the github repository

1 papers0 benchmarks

TaRBench (Test Case Repair Benchmark)

TaRBench is a comprehensive benchmark that we developed to evaluate the effectiveness of TaRGet in automated test case repair. The benchmark encompasses 45,373 broken test repairs across 59 open-source projects, providing a diverse and extensive dataset for assessing the capabilities of TaRGet.

1 papers0 benchmarks
PreviousPage 536 of 1000Next