TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

MMTB (Multi-Mission Tool Bench)

Our test data has undergone five rounds of manual inspection and correction by five senior algorithm researcher with years of experience in NLP, CV, and LLM, taking about one month in total. It boasts extremely high quality and accuracy, with a tight connection between multiple rounds of missions, increasing difficulty, no unusable invalid data, and complete consistency with human distribution. Its evaluation results and conclusions are of great reference value for subsequent optimization in the Agent direction.

1 papers0 benchmarks

COFFE (COFFE: A Code Efficiency Benchmark for Code Generation)

COFFE COFFE is a Python benchmark for evaluating the time efficiency of LLM-generated code. It is released by the FSE'25 paper "COFFE: A Code Efficiency Benchmark for Code Generation". You can also refer to the project webpage for more details.

1 papers0 benchmarksTexts

ViLCo (ViLCo-Bench)

We propose the first standardized benchmark in multimodal continual learning for video data, defining protocols for training and metrics for evaluation. This standardized framework allows researchers to effectively compare models, driving advancements in AI systems that can continuously learn from diverse data sources.

1 papers0 benchmarksImages, Texts, Videos

Vis-CheBI20

Molecules represent tokens of the language of chemistry, which underlies not only chemistry itself, but also scientific fields that use chemical information such as pharmacy, material science, and molecular biology. Existing molecular information is distributed across text books, publications, and patents. To describe structural information (spatial arrangement of atoms), molecules are commonly drawn as 2D images in such documents, which makes Optical Chemical Structure Understanding (OCSU) play an important role in molecule-centric scientific discovery.

1 papers0 benchmarksImages, Texts

FABA-Instruct

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksActions, Images, Texts

A Ball Collision Dataset (ABCD)

A Ball-Collision Dataset (ABCD) serves as a comprehensive benchmark for investigating the interaction dynamics of moving objects within 3D environments. It includes multimodal recordings of ball trajectories, captured under various conditions, including different elevation angles, flight lengths, and speeds. This dataset contains raw event, rgb and IMU data collected from an FPGA-based drone and 3D motion capture data of the drone (static) and a moving ball.

1 papers1 benchmarksImages, Videos

Haystack

Contrary to prior scene graph datasets, Haystack contains explicit negative annotations, i.e. annotations that a given relation does not have a certain predicate class. Negative annotations are helpful especially in the field of scene graph generation and open up a whole new set of possibilities to improve current scene graph generation models. Haystack is 100% compatible with existing panoptic scene graph datasets and can easily be integrated with existing evaluation pipelines.

1 papers0 benchmarks

Usage-related Questions

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksTexts

Gaze-CIFAR-10

We construct Gaze-CIFAR-10, a gaze-augmented image dataset based on the standard CIFAR-10 benchmark, enhanced with human eye-tracking annotations collected using the HTC VIVE Pro Eye headset. The original CIFAR-10 dataset consists of 60,000 color images across 10 categories, each with a resolution of $32 \times 32$ pixels. To enable reliable human gaze tracking, all images are upsampled to $1024 \times 1024$ using the Real-ESRGAN model.

1 papers1 benchmarksImages, Time series, Tracking

MERGE SPCS

This dataset contains pre-processed versions of datasets introduced in prior works. Additionally, it also contains new data that are pertinent to the paper.

1 papers0 benchmarksBiology, Biomedical, Images, Medical, Tables, Tabular

Baidu

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

SVOX Night

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

SVOX Sun

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

Tokyo 24/7

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

MSLS Val

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

Data Storage System Performance

IOPS and Latency measurements of a real data storage system

1 papers0 benchmarksTables, Tabular

Off-Topic

This dataset consists of synthetic LLM system prompts paired with user prompts, classified as either off-topic or on-topic. The aim is to provide realistic, real-world-inspired examples reflecting how large language models (LLMs) are used today for both open-ended and closed-ended tasks, such as text generation and classification. This dataset can be used for training and benchmarking off-topic guardrails.

1 papers0 benchmarks

LLM evaluation scores (Scores given by LLM according to preassigned score)

Dataset is a CSV file, that contains evaluation scores given by a panel of LLMs to responses produced by other LLMs . Responses regard a forecasting task assigned to multiple LLMs. The evaluation of the individual forecasts are performed according to 9 criteria indicated in the prompt. (see for details https://arxiv.org/abs/2412.09385).

1 papers0 benchmarksTabular

CropCOCO

CropCOCO is a validation-only dataset of COCO val 2017 images cropped such that some keypoints annotations are outside of the image. It can be used for keypoint detection, out-of-image keypoint detection and localization, person detection and amodal person detection.

1 papers0 benchmarksImages

IRBFD

We establish the first large benchmark called IRBFD to facilitate the research in the area of nonuniformity correction and infrared UAV target detection, which consists of 50,000 manually labeled infrared images with various nonuniformity levels, multi-scale UAV targets and rich backgrounds with target annotations.

1 papers0 benchmarksImages
PreviousPage 550 of 1000Next