TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Wearanize+ Dataset (v1.0)

Wearanize+ includes overnight sleep data from 130 participants (one night each) using three different wearable devices: Zmax headband, Empatica E4 wristband, and ActivPAL leg patch, alongside full-scale PSG recorded with SomnoScreen Plus and Mentalab Explore Pro. It also includes questionnaires, such as PSQI, MADRE, and PHQ-9, providing information on participants’ sleep, dreams, and overall health. (The link to access the dataset will be added soon).

1 papers0 benchmarksBiomedical, Tabular, Time series

Multi Lingual Bug Reports

Dataset Description The dataset used in this study comprises bug reports extracted from the Visual Studio Code GitHub repository, specifically focusing on those labeled with the english-please tag. This label indicates that the original submission was written in a language other than English, providing a clear signal for multilingual content. The dataset spans a five-year period (March 2019--June 2024), ensuring a diverse representation of bug types, user environments, and technical contexts.

1 papers1 benchmarksGraphs, Images, Texts

RawNIND (Raw Natural Image Noise Dataset)

The Raw Natural Image Noise Dataset (RawNIND) is a diverse collection of paired raw images designed to support the development of denoising models that generalize across sensors, image development workflows, and styles.

1 papers0 benchmarks

Bitcoin Historical Events Time-line 2009-2024

It includes 227 impactful events in Bitcoin history that shook the global markets. It can help evaluate Bitcoin market forecasting models in turbulent or event-full times.

1 papers0 benchmarks

LeetCode-Contest

Contains 80 questions of LeetCode weekly and bi-weekly contests released after March 2024.

1 papers0 benchmarks

R1-Onevision

The R1-Onevision dataset is a meticulously crafted resource designed to empower models with advanced multimodal reasoning capabilities. Aimed at bridging the gap between visual and textual understanding, this dataset provides rich, context-aware reasoning tasks across diverse domains, including natural scenes, science, mathematical problems, OCR-based content, and complex charts.

1 papers0 benchmarksImages, Texts

House Plant Species Dataset

Description:

1 papers0 benchmarks

VisCon-1M

VisCon-100K is a dataset specially designed to facilitate fine-tuning of vision-language models (VLMs) by leveraging interleaved image-text web documents. Derived from 45K web documents of the OBELICS dataset, this release contains 100K image conversation samples. GPT-4V is used to generate image-contextual captions, while OpenChat 3.5 converts these captions into diverse free-form and multiple-choice Q&A pairs. This approach not only focuses on fine-grained visual content but also incorporates the accompanying web context to yield superior performance. Using the same pipeline, but substituting our trained contextual captioner for GPT-4V, we also release the larger VisCon-1M dataset

1 papers0 benchmarksImages, Texts

VisCon-100K

VisCon-100K is a dataset specially designed to facilitate fine-tuning of vision-language models (VLMs) by leveraging interleaved image-text web documents. Derived from 45K web documents of the OBELICS dataset, this release contains 100K image conversation samples. GPT-4V is used to generate image-contextual captions, while OpenChat 3.5 converts these captions into diverse free-form and multiple-choice Q&A pairs. This approach not only focuses on fine-grained visual content but also incorporates the accompanying web context to yield superior performance. Using the same pipeline, but substituting our trained contextual captioner for GPT-4V, we also release the larger VisCon-1M dataset

1 papers0 benchmarksImages, Texts

Branched Deformable Linear Objects (BDLOs) Dataset

For each BDLO, dynamic trajectory data is captured in real-world settings using a motion capture system operating at 100 Hz when robots grasp the BDLO’s ends. For details on dataset usage, please refer to DEFT_train.py. For BDLO 1 and BDLO 3, we record dynamic trajectory data when one robot grasps the middle of the BDLO while the other robot grasps one of its ends.

1 papers0 benchmarks

FormalSpecCpp

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

IHEval (Evaluation on Instruction Hierarchy)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

TwinSynths

The TwinSynths dataset is a novel benchmark designed to overcome common limitations found in earlier synthetic image datasets, such as low image quality, inadequate content preservation, and limited class diversity. TwinSynths generates pairs of images where each synthetic image is visually identical to its real counterpart, ensuring that the essential content remains intact while showcasing the unique architectural features of the generative models used. TwinSynths comprises two subsets:

1 papers0 benchmarksImages

GENEUTRAL

Dataset Card for Dataset Name <!-- Provide a quick summary of the dataset. -->

1 papers0 benchmarksTexts

GENTER (GEnder Name TEmplates with pRonouns)

This dataset consists of template sentences associating first names ([NAME]) with third-person singular pronouns ([PRONOUN]), e.g., [NAME] asked , not sounding as if [PRONOUN] cared about the answer . after all , [NAME] was the same as [PRONOUN] 'd always been . there were moments when [NAME] was soft , when [PRONOUN] seemed more like the person [PRONOUN] had been .

1 papers0 benchmarksTexts

NAMEXTEND

This dataset extends NAMEXACT by including words that can be used as names, but may not exclusively be used as names in every context.

1 papers0 benchmarksTexts

NAMEXACT

This dataset contains names that are exclusively associated with a single gender and that have no ambiguous meanings, therefore being exact with respect to both gender and meaning.

1 papers0 benchmarksTexts

GENTYPES (Gender Stereotypes)

This dataset contains short sentences linking a first name, represented by the template mask [NAME], to stereotypical associations.

1 papers0 benchmarksTexts

UAV-PDD2023

UAV-PDD2023: A benchmark dataset for pavement distress detection based on UAV images

1 papers1 benchmarks

JASMIN-CGN (JASMIN-spraakcorpus)

A corpus of about 115 hours of Dutch speech from juveniles, non-native speakers and seniors, consisting of read text and man-machine dialogues.

1 papers0 benchmarks
PreviousPage 542 of 1000Next