TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Remote Flash LiDAR Vehicles Dataset

This dataset includes 3D point-cloud and 2D imagery from a flash LiDAR...

1 papers6 benchmarks3D, Images, LiDAR, Point cloud, Videos

Toxicity (Toxicity - UCI Machine Learning Repository)

Classification dataset of 171 molecules according to its toxicity from the UCI Machine Learning Repository.

1 papers0 benchmarks

Period Changer (Period Changer - UCI Machine Learning Repository)

Classification dataset of 90 molecules according to its effects in the circadian rhythm from the UCI Machine Learning Repository.

1 papers0 benchmarks

Federal and State-Level Election Results since 1955

DOI: https://doi.org/10.7910/DVN/O4CRXK

1 papers0 benchmarksTabular

MNIST+U

MNIST dataset with included uncertainty. For details see the corresponding paper.

1 papers0 benchmarks

RB-Dust (RB-Dust: Real-world Industrial Dust Dehazing Dataset)

A small-scale real-world dataset containing hazy/dusty industrial images and their clean ground truth counterparts. Designed for evaluating deep learning models for dust removal and image dehazing in industrial environments. Collected and fine-tuned by Moshtaghioun et al., 2025.

1 papers2 benchmarksImages

Metabolic Syndrome Dataset

The dataset utilized in this study consists of demographic, clinical, and laboratory data for 2,401 individuals. It includes 15 columns: 13 features, a sequence ID, and a response variable indicating the presence or absence of metabolic syndrome. The 13 features encompass age, sex, marital status, income, race, waist circumference, body mass index (BMI), albuminuria (presence of the blood protein albumin in urine), urine albumin-to-creatinine ratio, uric acid, blood glucose, high-density lipoprotein (HDL), and triglycerides. This dataset originates from the National Health and Nutrition Examination Survey (NHANES) database.

1 papers0 benchmarks

tibetan_news_classification

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

CADEdgeTune

CAD-EdgeTune dataset is acquired using a Husarion ROSbot 2.0 and ROSbot 2.0 Pro with the collection speed set to 5 frames per second from a suburban university environment. We may split the information into subgroups for noon, dusk, and dawn in order to depict our surroundings under various lighting situations. We have assembled 17 sequences totaling 8080 frames, of which 1619 have been manually analyzed using an open-source pixel annotation program. Since nearby photographs are highly similar to one another, we decide to annotate every five images. Since the annotation procedure may be highly time-consuming, we employ soft-labeling while annotating CAD-EdgeTune dataset, which enables us to proceed through the frames more quickly. The annotation method we employ enables us to create minute annotations inside an image's objects, and the categorization would encompass related pixels. This approach may result in less-than-perfect annotations and some performance accuracy loss, but the los

1 papers0 benchmarksImages

3 Sources (3 Sources Dataset)

We provide here a new multi-view text dataset, collected from three well-known online news sources: BBC, Reuters, and The Guardian. This dataset exhibits a number of common aspects of multi-view problems highlighted previously -- notably that certain stories will not be reported by all three sources (i.e incomplete views), and the related issue that sources vary in their coverage of certain topics (i.e. partially missing patterns).

1 papers0 benchmarksTexts

AutomotiveUI-Bench-4K

Dataset Overview: 998 images and 4,208 annotations focusing on interaction with in-vehicle infotainment (IVI) systems. Key Features:

1 papers0 benchmarksImages, Texts

SAS-Bench

SAS-Bench represents the first specialized benchmark for evaluating Large Language Models (LLMs) on Short Answer Scoring (SAS) tasks. Utilizing authentic questions from China's National College Entrance Examination (Gaokao), our benchmark offers:

1 papers0 benchmarks

AC-Bench

AC-Bench: A Benchmark for Actual Causality Reasoning Dataset Description AC-Bench is designed to evaluate the actual causality (AC) reasoning capabilities of large language models (LLMs). It contains a collection of carefully annotated samples, each consisting of a story, a query related to actual causation, detailed reasoning steps, and a binary answer. The dataset aims to provide a comprehensive benchmark for assessing the ability of LLMs to perform formal and interpretable AC reasoning.

1 papers0 benchmarksTexts

LeafNet (LeafNet: A large-scale dataset for training image-text models in leaf disease identification)

The PlantVillage dataset, with over 54,000 images spanning 14 plant species and 26 disease types, has been widely used for leaf disease classification. However, it is limited in both scale and diversity. To address these limitations, we developed LeafNet, a large-scale dataset designed to support foundation models for leaf disease diagnosis. LeafNet comprises over 186,000 images from 22 crop species, covering 43 fungal diseases, 8 bacterial diseases, 2 mould (oomycete) diseases, 6 viral diseases, and 3 mite-induced diseases, categorized into 97 classes. The dataset was meticulously collected and processed to minimize intra-class variations while ensuring clarity by maintaining a consistent imaging distance. The disease symptom descriptions were curated from reputable sources, including UME, NIH, and published studies, providing high-quality annotations to support AI-driven plant pathology research.

1 papers1 benchmarks

FLUXSynID

The FLUXSynID Dataset comprises 14,889 high-resolution synthetic face identities, each uniquely represented with paired images: document-style images and trusted live-capture images. The document-style images are designed to simulate ID photos, while the live-capture images are generated to reflect real-world conditions, including variations in lighting, facial expressions, and camera angles. This dual representation facilitates robust benchmarking for face recognition, morphing attack detection, and cross-domain identity verification tasks. FLUXSynID leverages cutting-edge generative models to ensure high fidelity and diversity across identities, enabling comprehensive analysis of biometric systems under varied conditions.

1 papers0 benchmarks

StoryReasoning

Visual storytelling systems struggle to maintain character identity across frames and link actions to appropriate subjects, frequently leading to referential hallucinations. These issues can be addressed through grounding of characters, objects, and other entities on the visual elements. We propose StoryReasoning, a dataset containing 4,178 stories derived from 52,016 movie images, with both structured scene analyses and grounded stories. Each story maintains character and object consistency across frames while explicitly modeling multi-frame relationships through structured tabular representations. Our approach features cross-frame object re-identification using visual similarity and face recognition, chain-of-thought reasoning for explicit narrative modeling, and a grounding scheme that links textual elements to visual entities across multiple frames. We establish baseline performance by fine-tuning Qwen2.5-VL 7B, creating Qwen Storyteller, which performs end-to-end object detection,

1 papers0 benchmarks

bespokelabs/Bespoke-Stratos-17k

Bespoke-Stratos-17k We replicated and improved the Berkeley Sky-T1 data pipeline using SFT distillation data from DeepSeek-R1 to create Bespoke-Stratos-17k -- a reasoning dataset of questions, reasoning traces, and answers.

1 papers0 benchmarks

NuminaMath-TIR (AI-MO/NuminaMath-7B-TIR)

Model Card for NuminaMath 7B TIR NuminaMath is a series of language models that are trained to solve math problems using tool-integrated reasoning (TIR). NuminaMath 7B TIR won the first progress prize of the AI Math Olympiad (AIMO), with a score of 29/50 on the public and private tests sets.

1 papers0 benchmarks

SynthCheX-75K

A dataset consisting of high-quality, synthetic chest X-rays from the CheXGenBench-benchmark leading model, Sana (0.6B). The dataset has been filtered to contain using High Quality samples using HealthGPT.

1 papers0 benchmarksImages

VADD (Video Anomaly Detection Dataset (VADD))

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers3 benchmarks
PreviousPage 555 of 1000Next