TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

BD-TypoSAT (Building Damage Typology Satellite Dataset)

On Sunday, August 29, 2021, Hurricane Ida struck parts of Louisiana and Mississippi with wind gusts reaching up to 172 mph, leaving more than a million customers without electricity, including the entire New Orleans area. During the disaster, Maxar captured high spatial resolution satellite imagery (at 0.4 m/pixel) and was subsequently made publicly available. The original images were segmented into 512*512-pixel patches to maintain spatial context while enabling detailed analysis. From this process, we generated a dataset of 2,135 triplets, each containing a pre-disaster image, a post-disaster image, and a manually annotated damage categorical mask.

1 papers0 benchmarksImages

EMMOE-100

first everyday task dataset featuring COT outputs, diverse task designs, detailed re-plan processes, along with SFT and DPO sub-datasets.

1 papers0 benchmarksImages, Texts

Experiments Dataset for PerfCam: Digital Twinning for Production Lines Using 3D Gaussian Splatting and Vision Models

This repository presents the dataset used in the PerfCam's original paper. This dataset is to support further research in the area of industrial 3D reconstruction, digital twinning, and predictive maintenance.

1 papers0 benchmarks

Finetune-RAG

This dataset is part of the Finetune-RAG project, which aims to tackle hallucination in retrieval-augmented LLMs. It consists of synthetically curated and processed RAG documents that can be utilised for LLM fine-tuning.

1 papers0 benchmarks

DLO Instance Segmentation dataset (DLO Instance Segmentation dataset generated by Blender)

Contains ~60000 HD images of Deformable Linear Objects (DLOs) generated using blender. The dataset contains a variety of industrial-looking backgrounds and contains instance segmentation masks. The main task for this dataset is DLO instance segmentation.

1 papers0 benchmarksImages

HAVOC (Harmful Abstractions and Violations in Open Completions Benchmark)

measure the toxicity generated by language models across input severity and harm categories, by creating a new benchmark of open ended prefixes. We sampled 10,376 snippets from web pages across the dimensions and harms as described in the paper - https://arxiv.org/pdf/2505.02009.

1 papers0 benchmarksTexts

MatTools

pymatgen_code_qa benchmark: qa_benchmark/generated_qa/generation_results_code.json, which consists of 34,621 QA pairs. pymatgen_code_doc benchmark: qa_benchmark/generated_qa/generation_results_doc.json, which consists of 34,604 QA pairs. real-world tool-usage benchmark: src/question_segments, which consists of 49 questions (138 tasks). One subfolder means a question with problem statement, property list and verification code.

1 papers0 benchmarksTexts

SOMPT22 (Surveillance Oriented Multi-Pedestrian Tracking Dataset (SOMPT22))

SOMPT22 is a multi-object tracking (MOT) benchmark focused on surveillance-style pedestrian tracking.

1 papers0 benchmarksImages, Tracking, Videos

GenoAdv

Genomics Adversarial Attack Sample dataset

1 papers0 benchmarksTexts

BlurRF-Synth

The first large-scale dataset for training and evaluating novel-view synthesis from blurred images.

1 papers0 benchmarks3D, Images

BlurRF-Real

A real-world low-light camera motion blur dataset for evaluating deblurring radiance fields methods.

1 papers0 benchmarks3D, Images

migration-bench-java-full

🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages.

1 papers0 benchmarksTexts

migration-bench-java-selected

🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages.

1 papers0 benchmarksTexts

migration-bench-java-utg

🤗 MigrationBench is a large-scale code migration benchmark dataset at the repository level, across multiple programming languages.

1 papers0 benchmarksTexts

STATE ToxiCN

With the rise of social media, user-generated content has surged, and hate speech has proliferated. Hate speech targets groups or individuals based on race, religion, gender, region, sexual orientation, or physical traits, expressing malice or inciting harm. Recognized as a growing social issue, it affects 9.41 billion Mandarin Chinese speakers (12% of the global population). However, research on Chinese hate speech detection lags, facing two key challenges.

1 papers0 benchmarksTexts

Bc8 (Bc8BioRED)

Bc8BioRED is built upon BioRED 2022 with the addition of directionality annotations. The training and development sets from the original 2022 BioRED corpus were combined and reused as the training set, while the test set was used as the development set. Furthermore, the 400 test abstracts from the BioCreative VIII were utilized for evaluation. Bc8BioRED encompasses seven types of entities and eight types of relationships. Each relationship annotation in the Bc8BioRED corpus is categorized by novelty to indicate whether the relationship represents a significant finding or previously known background knowledge. Initially, the BioRED 2022 corpus comprised 600 abstracts for RE system development, with an additional 400 abstracts annotated to enhance coverage of emerging topics. The dataset encompasses directionality annotations (subject/object roles) for each relation pair, resulting in 10,864 directionality annotations.

1 papers2 benchmarks

TQBA++ (Tiny QA Benchmark++)

Ultra-lightweight, multilingual QA eval dataset for rapid testing LLMs.

1 papers0 benchmarksTexts

Verireason-RTL-Coder_7b_reasoning_tb (VeriReason Verilog Dataset with Reasoning, Testbench, and Simulation Results)

Verireason-RTL-Coder_7b_reasoning_tb For implementation details, visit our GitHub repository: VeriReason

1 papers0 benchmarksTexts

Verireason-RTL-Coder_7b_reasoning_tb_simple (Simple Problems of VeriReason Verilog Dataset with Reasoning, Testbench, and Simulation Results)

Verireason-RTL-Coder_7b_reasoning_tb_simple For implementation details, visit our GitHub repository: VeriReason

1 papers0 benchmarksTexts

gen-reg

Node-level tasks.

1 papers0 benchmarks
PreviousPage 556 of 1000Next