TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

ConsisID-preview-Data

Description

1 papers0 benchmarksTexts, Videos

ChronoMagic-ProH

Description

1 papers0 benchmarksTexts, Videos

PDFM Embeddings (Population Dynamics Foundation Model Embeddings)

PDFM Embeddings are condensed vector representations designed to encapsulate the complex, multidimensional interactions among human behaviors, environmental factors, and local contexts at specific locations. These embeddings capture patterns in aggregated data such as search trends, busyness trends, and environmental conditions (maps, air quality, temperature), providing a rich, location-specific snapshot of how populations engage with their surroundings. Aggregated over space and time, these embeddings ensure privacy while enabling nuanced spatial analysis and prediction for applications ranging from public health to socioeconomic modeling.

1 papers0 benchmarksTabular

INCLUDE-base (44 languages)

Dataset Summary INCLUDE is a comprehensive knowledge- and reasoning-centric benchmark across 44 languages that evaluates multilingual LLMs for performance in the actual language environments where they would be deployed. It contains 22,637 4-option multiple-choice-questions (MCQ) extracted from academic and professional exams, covering 57 topics, including regional knowledge.

1 papers0 benchmarks

SBS Figures

Building a large-scale figure QA dataset requires a considerable amount of work, from gathering and selecting figures to extracting attributes like text, numbers, and colors, and generating QAs. Although recent developments in LLMs have led to efforts to synthesize figures, most of these focus primarily on QA generation. Additionally, creating figures directly using LLMs often encounters issues such as code errors, similar-looking figures, and repetitive content in figures. To address this issue, we present SBSFigures (Stage-by-Stage Synthetic Figures), a dataset for pre-training figure QA. Our proposed pipeline enables the creation of chart figures with complete annotations of the visualized data and dense QA annotations without any manual annotation process. Our stage-by-stage pipeline makes it possible to create diverse topic and appearance figures efficiently while minimizing code errors. Our SBSFigures demonstrate a strong pre-training effect, making it possible to achieve efficie

1 papers0 benchmarks

Student's EEG Brain Signal

This dataset consists of EEG (Electroencephalogram) recordings collected from students at our college during an educational experiment. The objective of this dataset is to evaluate students' cognitive engagement and learning effectiveness while interacting with educational content.

1 papers0 benchmarksTabular

MMComposition

MMCOMPOSITION is a high-quality benchmark specifically designed to comprehensively evaluate the compositionality of pre-trained Vision-Language Models (VLMs) across three main dimensions—VL compositional perception, reasoning, and probing—which are further divided into 13 distinct categories of questions. While previous benchmarks have mainly focused on text-to-image retrieval, single-choice questions, and open-ended text generation, MMCOMPOSITION introduces a more diverse and challenging set of 4,342 tasks covering both single-image and multi-image scenarios, as well as single-choice and indefinite-choice formats. This expanded range of tasks aims to capture the complex interplay between vision and language more effectively, surpassing earlier benchmarks such as ARO and Winoground by providing a more comprehensive and in-depth assessment of models’ cross-modal compositional capabilities.

1 papers0 benchmarksImages, Texts

Financial Dynamic Knowledge Graph

FinDKG: The Global Financial Dynamic Knowledge Graph Dataset FinDKG is an open-source dataset focused on creating a temporally-resolved Financial Dynamic Knowledge Graph. Designed to bridge the gap in industry-specific knowledge graphs, particularly in the financial sector, FinDKG provides a high-touch, temporally-aware representation of global economic and market dynamics. This repository includes comprehensive details about the dataset, methodology, and schema, aiming to facilitate academic research and actionable insights in global financial markets.

1 papers0 benchmarksFinancial, Graphs, Texts

Plancraft

An evaluation dataset for planning with LLM agents

1 papers0 benchmarksEnvironment, Images, Texts

AirLetters

This large collection of over 161,000 video-label pairs of video clips, shows humans drawing letters and digits in the air, and is used to evaluate a model’s ability to classify articulated motions correctly. Unlike existing video datasets, AirLetters’ accurate classification predictions rely on discerning motion patterns and integrating information presented by the video over time (i.e., over many frames of video). That study revealed that while trivial for humans, accurate representations of complex articulated motions remain an open problem for end-to-end learning for video understanding models.

1 papers0 benchmarksVideos

Wiki-ImageReview1.0 (https://huggingface.co/datasets/naist-nlp/Wiki-ImageReview1.0)

We introduce a novel task for LVLMs, which involves reviewing the good and bad points of a given image. We construct a benchmark dataset containing 207 images selected from Wikipedia. Each image is accompanied by five review texts and a manually annotated ranking of these texts in both English and Japanese.

1 papers0 benchmarks

MapEval

MapEval contains 700 question-answer pairs.

1 papers0 benchmarksTexts

MapEval-Textual

MapEval-Textual contains 300 context-question-answer triplets. The necessary geo-spatial information is provided in the context. The task is to answer question based on the factual data provided in the context.

1 papers1 benchmarksTexts

MapEval-Visual

MapEval-Visual contains 400 image-question-answer triplets. Each question is paired with a snapshot from google maps website. The task is the answer question based on the provided map snapshot.

1 papers2 benchmarksImages, Texts

Persuasive Writing Strategy

Persuasive Writing Strategy dataset on the health subset of the Multi-FC dataset.

1 papers0 benchmarks

MapEval-API

MapEval-Textual contains 300 question-answer pairs. The task is to answer question by fetching necessary informations using external Map APIs.

1 papers1 benchmarksTexts

NYUDv2-IS

A RGB-D dataset converted from NYUDv2 into COCO-style instance segmentation format. To construct NYUDv2-IS, specifically tailored for instance segmentation, we generated instance masks that delineate individual objects in each image. These masks were labeled using the object class annotations provided in the original NYUDv2 dataset, which is distributed in MATLAB format. The process involved several key steps: (1) extracting binary instance masks, (2) converting these masks into polygon representations, and (3) generating COCO-style annotations. Each annotation includes essential attributes such as category ID, segmentation masks, bounding boxes, object areas, and image metadata. During this conversion, we focused on 9 categories out of the original 13 classes, excluding non-instance categories such as walls and floors. To ensure dataset quality, images without any object annotations were systematically removed.

1 papers1 benchmarksImages, RGB-D

SUN-RGBD-IS

A RGB-D dataset converted from SUN-RGBD into COCO-style instance segmentation format. To transform SUN-RGBD into an instance segmentation benchmark (i.e., SUN-RGBDIS), we employed a pipeline similar to that of NYUDv2-IS. We selected 17 categories from the original 37 classes, carefully omitting non-instance categories like ceilings and walls. Images lacking any identifiable object instances were filtered out to maintain dataset relevance for instance segmentation tasks. We systematically convert segmentation annotations into COCO format, generating precise bounding boxes, instance masks, and object attributes.

1 papers1 benchmarksImages, RGB-D

Box-IS

RGB-D instance segmentation box dataset. The Box-IS dataset was created to support research on human-robot collaboration with a focus on robotic manipulation tasks. It was captured using the Intel® RealSense™ Depth Camera D455, a high-performance sensor designed for depth imaging. To ensure precise depth measurements, we bypassed the default depth data processing of the sensor and performed accurate stereo matching directly from the captured left and right IR images. Employing the UniMatch technique, we derived a high-quality depth map from these stereo IR images, which was then aligned with the corresponding RGB image for a comprehensive output. The dataset was intentionally designed to encompass a broad range of scene complexities, from simple box arrangements to highly irregular configurations. This diversity ensures that it can effectively benchmark algorithms across varying levels of difficulty.

1 papers1 benchmarksImages, RGB-D

JoyGen-TFace

A high-definition Talking Face dataset featuring Chinese-language videos

1 papers0 benchmarks
PreviousPage 535 of 1000Next