TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

SMC

A. Holzapfel, M. E. Davies, J. R. Zapata, J. L. Oliveira, and F. Gouyon, “Selective sampling for beat tracking evaluation,” Transactions on Audio, Speech, and Language Processing, vol. 20, no. 9, pp. 2539–2548, 2012

1 papers1 benchmarksAudio

TapCorrect

J. Driedger, H. Schreiber, W. B. de Haas, and M. Müller, “Towards automatically correcting tapped beat annotations for music recordings.” in Proceedings of the International Society for Music Information Retrieval Conference (ISMIR), 2019

1 papers2 benchmarksAudio

WikiOFGraph (Wikipedia Ontology-Free Graph-Text)

a high-level explanation of the dataset characteristics We introduce WikiOFGraph, a novel large-scale, domain-diverse dataset synthesized by LLMs, ensuring superior graph-text consistency to advance general-domain graph-to-text generation.

1 papers2 benchmarksGraphs, Texts

ViCoS Towel Dataset

The ViCoS Towel Dataset is a state-of-the-art benchmark for grasp point localization on cloth objects, specifically towels. Designed to advance research in robotic grasping and perception for textile objects, this dataset includes a collection of 8,000 high-resolution RGB-D images (1920×1080) captured with a Kinect V2 under a variety of conditions. Each image provides detailed depth information, making it ideal for training deep learning models and conducting thorough benchmarking.

1 papers3 benchmarksImages, RGB-D

DreamVoiceDB

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksSpeech

Deformable Linear Objects (DLOs) Dataset

For each DLO, we collect 350 seconds of dynamic trajectory data in the real-world using the motion capture system at a frequency of 100 Hz.

1 papers0 benchmarks

AceParse

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksImages, Texts

RSM-based multi-objective optimization using desirability functions

The following files contains the simulation inputs and outputs for conducting the multi-objetive optimization of thermal comfort and dyalight with the Response Surface Methodology. This files feed are needed for running the R script/code as well as the datasets are contained in the Github repository.

1 papers0 benchmarksTabular

FDG-PET-CT-Lesions | A whole-body FDG-PET/CT dataset with manually annotated tumor lesions

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

JAMBO (A Multi-Annotator Image Dataset for Benthic Habitat Classification)

The JAMBO dataset contains 3290 underwater images of the seabed captured by an ROV in temperate waters in the Jammer Bay area off the North West coast of Jutland, Denmark. All the images have been annotated by six annotators to contain one of three classes: sand, stone, or bad.

1 papers0 benchmarksEnvironment, Images

ModSec–Learn dataset

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

DailyMoth-70h

DailyMoth-70h is a fully self-contained ASL-to-English sign language dataset containing over 70h of video (48K clips) with aligned English captions of a single native ASL signer (white, male, and early middle-aged) from the ASL news channel TheDailyMoth. The primary purpose of the dataset is to be used as a benchmark and analysis dataset for (gloss-free) sign language translation.

1 papers0 benchmarksTexts, Videos

SBA (Sequentail Brick Assembly Dataset)

The RAD (Randomly Assembled Object Construction) dataset is a synthetic 3D LEGO dataset designed for the task of Sequential Brick Assembly (SBA). Here are the key characteristics and details:

1 papers0 benchmarks3D, 3d meshes, Actions, Images

GerMS-AT (GERMS-AT: A Sexism/Misogyny Dataset of Forum Comments from an Austrian Online Newspaper)

This dataset contains 7984 user comments from an Austrian online newspaper. The comments have been annotated by 4 or more out of 11 annotators as to how strong sexism/mysogyny is present in the comment. It was used in the GermEval 2024 Shared Task 1: GerMS-Detect to evaluate data-driven approaches to automatically detect sexism in user comments.

1 papers2 benchmarksTexts

Machine Learning for Analyzing Atomic Force Microscopy (AFM) Images Generated from Polymer Blends

Dataset used in "Machine Learning for Analyzing Atomic Force Microscopy (AFM) Images Generated from Polymer Blends". Contains AFM images in raw ".ibw" format.

1 papers0 benchmarks

aiMotive 3D Traffic Light and Traffic Sign Dataset

A large-scale traffic sign and traffic light dataset with accurate 3D positioning and temporally consistent 3D bounding boxes of traffic management objects from up to 200 meters away. The dataset contains additional attributes such as traffic light state, traffic light mask type, traffic sign type, and occlusion. The application areas are 3D traffic lights and sign detection for autonomous driving.

1 papers0 benchmarks3D, Images

Data on COVID-19 by Our World in Data

COVID-19 dataset for the world.

1 papers0 benchmarks

mdCATH (mdCATH: A Large-Scale MD Dataset for Data-Driven Computational Biophysics)

This dataset comprises all-atom systems for 5,398 CATH domains, modeled with a state-of-the-art classical force field, and simulated in five replicates each at five temperatures from 320 K to 450 K.

1 papers0 benchmarks

SCI (Self-Contradictory Instructions)

Large multimodal models (LMMs) excel in adhering to human instructions. However, self-contradictory instructions may arise due to the increasing trend of multimodal interaction and context length, which is challenging for language beginners and vulnerable populations. We introduce the Self-Contradictory Instructions benchmark to evaluate the capability of LMMs in recognizing conflicting commands. It comprises 20,000 conflicts, evenly distributed between language and vision paradigms. It is constructed by a novel automatic dataset creation framework, which expedites the process and enables us to encompass a wide range of instruction forms. Our comprehensive evaluation reveals current LMMs consistently struggle to identify multimodal instruction discordance due to a lack of self-awareness. Hence, we propose the Cognitive Awakening Prompting to inject cognition from external, largely enhancing dissonance detection.

1 papers0 benchmarksImages, Texts

SemTabNet

Dataset Card for SemTabNet This dataset accompanies the following paper:

1 papers1 benchmarksTables, Tabular, Texts
PreviousPage 518 of 1000Next