TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

U2-Bench

U2-BENCH is the first large-scale benchmark for evaluating Large Vision-Language Models (LVLMs) on ultrasound imaging understanding. It provides a diverse, multi-task dataset curated from 40 licensed sources, covering 15 anatomical regions and 8 clinically inspired tasks across classification, detection, regression, and text generation.

1 papers0 benchmarksImages, Medical

ProjectEval

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksTexts

CUTS

This is the dataset released along with the publication:

1 papers0 benchmarksImages

GlobalGeoTree

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

MIBot - Motivational Interviewing for Smoking Cessation Dataset, based on MIBOT Version 6.3A

The dataset from the study "A Fully Generative Motivational Interviewing Counsellor Chatbot for Moving Smokers Towards the Decision to Quit". The dataset comprises annotated transcripts and surveys (including self-reported readiness to quit smoking) from 106 conversations between human smokers and MIBot v6.3A — a motivational interviewing (MI) chatbot built using OpenAI's GPT-4o.

1 papers0 benchmarksTexts

BiasLab

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

military_vehicles

Military Vehicles for Hierarchical Multi-label Classification (HMC)

1 papers0 benchmarks

ASyMOB (ASyMOB: Algebraic Symbolic Mathematical Operations Benchmark)

ASyMOB (pronounced Asimov, in tribute to the renowned author), is a novel assessment framework focused exclusively on symbolic manipulation, featuring 17,092 unique math challenges, organized by similarity and complexity. ASyMOB enables analysis of LLM failure root-causes and generalization capabilities by comparing performance in problems that differ by simple numerical or symbolic "perturbations".

1 papers0 benchmarksTexts

Steerbench

Steerability probe example for text-rewriting.

1 papers0 benchmarks

GLARE: Google Apps Arabic Reviews Dataset

GLARE is an Arabic Apps Reviews dataset collected from Saudi Google PlayStore. It consists of 76M reviews, 69M of which are Arabic reviews from 9,980 Android Applications. We present the data collection methodology, along with a detailed Exploratory Data Analysis (EDA) and Feature Engineering on the gathered reviews. We also highlight possible use cases and benefits of the dataset.

1 papers0 benchmarksTexts

Laboratory effect perception during virtual stages auralization

Introduction

1 papers0 benchmarksAudio

TimeGraph (TimeGraph: Synthetic Benchmark Datasets for Robust Time-Series Causal Discovery)

TimeGraph is a comprehensive suite of synthetic datasets designed to benchmark causal discovery algorithms on time-series data. The dataset captures real-world complexities by incorporating temporal dynamics such as trends, seasonality, and nonstationarity, as well as sampling challenges including irregular time intervals and structured missingness. It features diverse noise types, including Gaussian, heavy-tailed, and heteroskedastic variations, and supports scenarios with latent confounding to enable evaluation under partially observed systems. The underlying causal structures span both linear and nonlinear relationships, including polynomial and trigonometric forms.

1 papers0 benchmarksTabular, Time series

PRMC_L2 (Rock, Punk, Metal, and Core - Livehouse Lighting)

Dataset for studying the relationship between music and lighting in live music performances

1 papers0 benchmarks

Interaction Dataset of Autonomous Vehicles with Traffic Lights and Signs (Interaction Data of Autonomous Vehicles with Traffic Lights and Signs Based on Waymo Motion Open Dataset)

This dataset is derived from the Waymo Motion dataset and focuses on capturing the interactions between autonomous vehicles (AVs) and traffic control devices such as traffic lights and stop signs. It addresses a critical gap by providing real-world trajectory data that reflects how AVs interpret and respond to traffic control signals, supporting research in AV behavior modeling, traffic simulation, and the design of intelligent transportation systems.

1 papers0 benchmarksTime series, Tracking

Upper body thermal images and associated clinical data from a pilot cohort study of COVID-19

The prospective upper body thermal images SARS-CoV2 association study was designed to test the hypothesis that thermal videos may aid in the early diagnosis of COVID-19. The study recorded a set of measurements from 252 participants regarding PCR results, demographics, vital signs, participant activities, medications, respiratory symptoms, and a thermal video session where the volunteers performed simple breath-hold in four different positions. The acquired data may be used to test clinical association questions regarding temperature patterns, demographics, and vital signs. Furthermore, it could be valuable to develop new computer algorithms for extracting useful scientific information from thermal videos.

1 papers0 benchmarksImages, Tabular

NoMusic

NoMusic - The Norwegian Multi-Dialectal Slot and Intent Detection Corpus https://aclanthology.org/2024.vardial-1.9/

1 papers0 benchmarks

LEMONADE

LEMONADE is a large, expert-annotated dataset for event extraction from news articles in 20 languages: English, Spanish, Arabic, French, Italian, Russian, German, Turkish, Burmese, Indonesian, Ukrainian, Korean, Portuguese, Dutch, Somali, Nepali, Chinese, Persian, Hebrew, and Japanese.

1 papers0 benchmarksTexts

B-XAIC

B-XAIC consists of 50K small molecules represented as graphs and includes 7 graph classification tasks, each with ground truth labels and corresponding explanations.

1 papers0 benchmarksGraphs

ExeCheck

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

CAGUI (Chinese Android GUI Benchmark)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksImages, Texts
PreviousPage 560 of 1000Next