TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

SAGC-A68 (A space access graph dataset for the classification of spaces and space elements in apartment buildings)

The analysis of building models for usable area, building safety, and energy efficiency requires accurate classification data of spaces and space elements. To reduce input model preparation effort and errors, automated classification of spaces and space elements is desirable. Although existing space function classifiers use space adjacency or connectivity graphs as input, the application of Graph Deep Learning (GDL) to space layout element classification has not been extensively researched due to the lack of suitable datasets. To bridge this gap, we introduce a dataset named SAGC-A68, which comprises access graphs automatically generated from 68 digital 3D models of space layouts of apartment buildings designed or built between 1952 and 2019 in 13 countries. Each access graph contains nodes representing spaces and space elements and edges representing the connection between them. Nodes are uniquely identified and characterized by 16 features including “Position X”, “Position Y”, “Posit

1 papers0 benchmarks

ConvSumX

ConvSumX is a cross-lingual conversation summarization benchmark, through a new annotation schema that explicitly considers source input context. ConvSumX consists of 2 sub-tasks under different real-world scenarios, with each covering 3 language directions.

1 papers0 benchmarksTexts

OCFR-LFW (Occluded Face Recognition LFW)

A occluded version of the LFW dataset for occluded face recognition verification. Uses structured occlusions generated to seem more realistic.

1 papers0 benchmarks

Multicenter dataset of simulated neuroimaging features - quadratic relationship with age

A detailed description of this dataset can be found in the Zenodo repository: https://zenodo.org/record/8119042#.ZK-jJC9BxhE

1 papers0 benchmarksTabular

Multicenter dataset of neuroimaging features (part I)

A detailed description of this dataset can be found in the Zenodo repository: https://zenodo.org/record/7845311#.ZK-jty9BxhE

1 papers0 benchmarksTabular

Multicenter dataset of neuroimaging features (part II)

A detailed description of this dataset can be found in the Zenodo repository: https://zenodo.org/record/7845361#.ZK-k7y9BxhE

1 papers0 benchmarksTabular

Subjective Perception of Active Noise Reduction (SPANR) (Replication Data for: Anti-noise window: subjective perception of active noise reduction and effect of informational masking)

This repository contains replication data to the paper titled: "Anti-noise window: subjective perception of active noise reduction and effect of informational masking"

1 papers0 benchmarksTexts

CBTex (Synthetic CardBoard Textures)

Dataset of >200 synthetic cardboard texture images that were rendered with DoubeGum's cardboard shader in Blender. Used to generate Parcel3D, the dataset for our paper on single image 3D reconstructions of potentially damaged parcels.

1 papers0 benchmarksImages

HabiCrowd

HabiCrowd, a new dataset and benchmark for crowd-aware visual navigation that surpasses other benchmarks in terms of human diversity and computational utilization. HabiCrowd can be utilized to study crowd-aware visual navigation tasks. A notable feature of HabiCrowd is that our crowd-aware settings is 3D, which is scarcely studied by previous works.

1 papers0 benchmarks

BDD-QA

BDD-QA is distinguished by its encompassing range of traffic actions, crafted to rigorously evaluate a model's decision-making abilities in traffic scenario. This makes it a potent tool for high-level decision-making research within traffic contexts, including autonomous driving developments.

1 papers0 benchmarks

HDT-QA (human driving test question answering dataset)

HDT-QA, coupled with driving manuals, offers an extensive compendium of driving instructions and driving knowledge tests across all 51 states of the US. This resource is beneficial for assessing the incorporation and impact of traffic knowledge within intelligent driving systems, marking a crucial stride towards more advanced, informed, and safe autonomous driving technology.

1 papers0 benchmarks

Complex-TV-QA (complex traffic video question answer datset)

The Complex-TV-QA dataset, to our knowledge, is the inaugural resource that provides human-annotated, detailed video captions within traffic scenarios, alongside complex reasoning questions. This novel dataset not only stands as a vital tool for evaluating language models in real-world video-QA and video-reasoning research, but also offers valuable insights for the development and understanding of multi-modal video reasoning models and related works.

1 papers0 benchmarks

NEU dataset

Data set used in the work One-Shot Recognition of Manufacturing Defects in Steel Surfaces

1 papers0 benchmarks

SHD - Adding (Spiking Heidelberg Digits - Adding)

This dataset is based on the Spiking Heidelberg Digits (SHD) dataset. Sample inputs consist of two spike encoded digits sampled uniformly at random from the SHD dataset and concatenated, with the target being the sum of the digits (irrespective of language). The train and test split remain the same, with the test set consisting of 16k such samples based on the SHD test set.

1 papers1 benchmarksAudio

WYWEB (https://github.com/baudzhou/WYWEB)

An evaluation bentchmark for classical Chinese.

1 papers0 benchmarks

RePoGen

Synthetic humans generated by the RePoGen method.

1 papers0 benchmarksImages

PatchDB

PatchDB is a large-scale security patch dataset that contains around 12K security patches and 24K non-security patches from the real world.

1 papers0 benchmarks

Zucker HRI Dataset

Zucker HRI Dataset contains two different agent types (robot and human) in several scenarios. The robot switched between 3 different motion controllers (Linear, NHTTC, and CADRL) over multiple different scenarios with different permutations of human agents. There are also scenes without the robot for a baseline.

1 papers0 benchmarks

COLLIE-v1

COLLIE-v1 is a dataset with 2080 instances comprising 13 constraint structures designed for text generation under constraints. It is a grammar-based framework that allows the specification of rich, compositional constraints with diverse generation levels (word, sentence, paragraph, passage).

1 papers0 benchmarksTexts

Experimental Results for "A Unified Perspective on Natural Gradient Variational Inference with Gaussian Mixture Models"

This package contains the raw data / logs (fetched from WandB) for the experiments of the following publication:

1 papers0 benchmarks
PreviousPage 466 of 1000Next