TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

RiskData

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksTexts

RASMD (RASMD: RGB And SWIR Multispectral Driving Dataset for Robust Perception in Adverse Conditions)

Current autonomous driving algorithms heavily rely on the visible spectrum, which is prone to performance degradation in adverse conditions like fog, rain, snow, glare, and high contrast. Although other spectral bands like near-infrared (NIR) and long-wave infrared (LWIR) can enhance vision perception in such situations, they have limitations and lack large-scale datasets and benchmarks. Short-wave infrared (SWIR) imaging offers several advantages over NIR and LWIR. However, no publicly available large-scale datasets currently incorporate SWIR data for autonomous driving. To address this gap, we introduce the RGB and SWIR Multispectral Driving (RASMD) dataset, which comprises 100,000 synchronized and spatially aligned RGB-SWIR image pairs collected across diverse locations, lighting, and weather conditions. In addition, we provide a subset for RGB-SWIR translation and object detection annotations for a subset of challenging traffic scenarios to demonstrate the utility of SWIR imaging t

1 papers0 benchmarksImages

GenoTEX (An LLM Agent Benchmark for Automated Gene Expression Data Analysis)

GenoTEX (Genomics Data Automatic Exploration Benchmark) is a benchmark dataset for the automated analysis of gene expression data to identify disease-associated genes while considering the influence of other biological factors. It provides analysis code and results for solving a wide range of gene-trait association (GTA) analysis problems, encompassing dataset selection, preprocessing, and statistical analysis, in a pipeline that follows computational genomics standards. The benchmark includes expert-curated annotations from bioinformaticians to ensure accuracy and reliability.

1 papers0 benchmarksTabular, Texts

taste-music-dataset (Taste Music Dataset)

This dataset is a patched version of The Taste & Affect Music Database by D. Guedes et al. It is a set of captions that describe 100 musical pieces and associate with them gustatory keywords on the basis of Guedes findings.

1 papers0 benchmarksAudio, Music, Texts

GraspClutter6D

GraspClutter6D is a large-scale real-world dataset for robust object perception and robotic grasping in cluttered environments. It features 1,000 highly cluttered scenes with dense arrangements (average 14.1 objects/scene with 62.6% occlusion), 200 household, industrial, and warehouse objects captured in 75 diverse environment configurations (bins, shelves, and tables), multi-view data from 4 RGB-D cameras (RealSense D415, D435, Azure Kinect, and Zivid One+), and comprehensive annotations including 736K 6D object poses and 9.3 billion feasible robotic grasps for 52K RGB-D images. The dataset provides a challenging testbed for segmentation, 6D pose estimation, and grasp detection algorithms in realistic cluttered scenarios.

1 papers0 benchmarks3d meshes, 6D, Images, RGB-D

tcm-llm-overrely-on-names

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

RoBo6

Dataset contains light curves of 6 rocket body types from Mini Mega Tortora database (MMT)[^1]. The dataset was created to be used as a benchmark for rocket body light curve classification. For more informations follow the original paper: RoBo6: Standardized MMT Light Curve Dataset for Rocket Body Classification[^2]

1 papers0 benchmarksTime series

DPLink-ISP-Shanghai

dataset in WWW 2019 "DPLink: User Identity Linkage via Deep Neural Network From Heterogeneous Mobility Data". This data is intended for academic use only. Redistribution of this data is not permitted without our explicit permission.

1 papers0 benchmarksTime series

Source code (Source code underlying the publication: Topology-Based Reconstruction Prevention for Decentralised Learning)

MATLAB code to reproduce results presented in the paper "Topology-Based Reconstruction Prevention for Decentralised Learning".

1 papers0 benchmarks

Major TOM Core-DEM

ML-ready Global Dataset of elevation map. Adapting Copernicus DEM GLO-30 to the Major TOM framework.

1 papers0 benchmarksRGB-D

Corporación Favorita Grocery Sales Forecasting

The dataset has thousands of time series. The provided code select 40 times series of four different profiles to compare classical models (ARIMA) and deep learning models (LSTM)

1 papers0 benchmarks

World Values Survey

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

VOCEdits

VOCEdits: A benchmark for precise geometric object-level editing Sample format: (input image, edit prompt, input mask, ground-truth output mask, ...)

1 papers0 benchmarks

GitBugs (GitBugs: Bug Reports for Duplicate Detection, Retrieval Augmented Generation, Triage, and More)

GitBugs is a comprehensive and up-to-date dataset comprising over 150,000 bug reports from nine actively maintained open-source projects, including Firefox, Cassandra, and VS Code. GitBugs aggregates data from Github, Bugzilla and Jira issue trackers, offering standardized categorical fields for classification tasks and predefined train/test splits for duplicate bug detection. In addition, it includes exploratory analysis notebooks and detailed project-level statistics, such as duplicate rates and resolution times. GitBugs supports various software engineering research tasks, including duplicate detection, retrieval augmented generation, resolution prediction, automated triaging, and temporal analysis. The openly licensed dataset provides a valuable cross-project resource for bench- marking and advancing automated bug report analysis. Access the data and code at this https URL.

1 papers0 benchmarks

DivShift-NAWC (DivShift - North American West Coast)

DivShift North American West Coast DivShift Paper | Extended Version | Code

1 papers0 benchmarksImages

TGB (Temporal Graph Benchmark)

TGB is a collection of challenging and diverse benchmark datasets for realistic, reproducible, and robust machine learning evaluation on temporal graphs. It includes dynamic link and node property prediction tasks and an automated pipeline from dataset downloading, data loading, evaluation, and submission to the TGB leaderboard. TGB 2.0 includes novel datasets for temporal knowledge graphs and temporal heterogeneous graphs.

1 papers0 benchmarks

Discrete-Time Modeling of Interturn Short Circuits in Interior PMSMs - Data and Models

Project: Discrete-Time Modeling of Interturn Short Circuits in Interior PMSMs

1 papers0 benchmarksTime series

CAShift (Cloud Attack & Normality Shift Dataset)

CAShift is the first multiple normality shift-aware Log-Based Anomaly Detection (LAD) dataset specifically designed for cloud systems, which considers different software roles in cloud systems and attack behavior among cloud components.

1 papers0 benchmarks

CCUP

a new self-annotated CC-ReID dataset named Cloth-Changing Unreal Person.

1 papers0 benchmarks

BlenderGym

BlenderGym is the first comprehensive VLM system benchmark for 3D graphics editing. It evaluates VLM systems through code-based 3D reconstruction tasks. BlenderGym consists of 245 hand-crafted Blender scenes across 5 key graphics editing tasks: procedural geometry editing, lighting adjustments, procedural material design, blend shape manipulation, and object placement.

1 papers0 benchmarks
PreviousPage 551 of 1000Next