TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Social-HM3D

The scene derives from photo-realistic HM3D datasets. Our dataset offers a wide variety of environments especially for Social Navigation tasks, with carefully calibrated human density, incorporating realistic human motions and natural movement patterns. These features ensure balanced interaction dynamics across diverse scenes, facilitating the development of more effective social navigation algorithms.

1 papers0 benchmarksEnvironment

Social-MP3D

The scene derives from photo-realistic MP3D datasets. Our dataset offers a wide variety of environments especially for Social Navigation tasks, with carefully calibrated human density, incorporating realistic human motions and natural movement patterns. These features ensure balanced interaction dynamics across diverse scenes, facilitating the development of more effective social navigation algorithms.

1 papers0 benchmarksEnvironment

NuiSI Dataset (Nuitrack Skeleton Interaction Dataset)

The NuiSI dataset contains skeleton tracking trajectories of Human Interaction Partners performing a variety of physically interactive behaviors (waving, handshaking, rocket fistbump, parachute fistbump) with each other. This is inspired by the dataset in Bütepage et al. "Imitating by generating: Deep generative models for imitation of interactive tasks." Frontiers in Robotics and AI (2020) wherein they capture a dataset with rokoko motion capture suits. Instead we track the skeletons of the interaction partner with Intel Realsense cameras using Nuitrack, for a more realistic scenario, with noise coming from the depth sensor, the skeleton tracking and some partial occlusions. This makes it more representative of real world interactions with a Robot equipped with an RGBD camera. T This dataset is used in our papers for training Interaction models for Human-Robot Interaction with a humanoid social robot. If you find the dataset useful in your work, please cite our paper:

1 papers0 benchmarks3D, Tracking

Human-Robot Interaction Data

Dataset used in Bütepage, Judith, et al. "Imitating by generating: Deep generative models for imitation of interactive tasks." Frontiers in Robotics and AI 7 (2020): 47.

1 papers0 benchmarks

Overnight

Overnight is a dataset for semantic parsing in eight domains.

1 papers0 benchmarks

Marxism small

This dataset is a collection of marxist fragments mixed and cut randomly from the Marxist archive (marxists.org).

1 papers0 benchmarks

SCARED-C (SCARED-Corrupted)

The dataset SCARED-C is introduced in the context of assessing robustness in endoscopic depth prediction models. It is part of the EndoDepth benchmark, which is designed to evaluate the performance of monocular depth prediction models specifically for endoscopic scenarios. The dataset features 16 different types of image corruptions, each with five levels of severity, encompassing challenges like lens distortion, resolution alterations, specular reflection, and color changes that are typical in endoscopic imaging. The ground truth is on the original testing set of SCARED.

1 papers2 benchmarksBiomedical, Images, Medical

MMInstruct-GPT4V (MMInstruct: A High-Quality Multi-Modal Instruction Tuning Dataset with Extensive Diversity)

Vision-language supervised fine-tuning effectively enhances VLLM performance, but existing visual instruction tuning datasets have limitations:

1 papers0 benchmarksImages, Texts

assamese-dataset (Assamese Dataset)

A collection of various NLP datasets in Assamese. These datasets are split into two: pre-training corpora and fine-tuning datasets.

1 papers0 benchmarks

Bukva (Bukva: Russian Sign Language Alphabet)

We introduce a video dataset Bukva for Russian Dactyl Recognition task. Bukva dataset size is about 27 GB, and it contains 3757 RGB videos with more than 101 samples for each RSL alphabet sign, including dynamic ones. The dataset is divided into training set and test set by subject user_id. The training set includes 3097 videos, and the test set includes 660 videos. The total video recording time is ~4 hours. About 17% of the videos are recorded in HD format, and 70% of the videos are in FullHD resolution.

1 papers1 benchmarksRGB Video, Videos

FairFD

A Racial Fairness Benchmark Dataset for Face Forgery Detection.

1 papers0 benchmarksImages

JamPatoisNLI

JamPatoisNLI provides the first dataset for natural language inference in a creole language, Jamaican Patois. Many of the most-spoken low-resource languages are creoles. These languages commonly have a lexicon derived from a major world language and a distinctive grammar reflecting the languages of the original speakers and the process of language birth by creolization. This gives them a distinctive place in exploring the effectiveness of transfer from large monolingual or multilingual pretrained models. While our work, along with previous work, shows that transfer from these models to low-resource languages that are unrelated to languages in their training set is not very effective, we would expect stronger results from transfer to creoles. Indeed, our experiments show considerably better results from few-shot learning of JamPatoisNLI than for such unrelated languages, and help us begin to understand how the unique relationship between creoles and their high-resource base languages af

1 papers1 benchmarks

RClicks

We conducted a large crowdsourcing study of click patterns in an interactive segmentation scenario and collected 475K real-user clicks. Drawing on ideas from saliency tasks, we develop a clickability model that enables sampling clicks, which closely resemble actual user inputs. Using our model and dataset, we propose RClicks benchmark for a comprehensive comparison of existing interactive segmentation methods on realistic clicks. Specifically, we evaluate not only the average quality of methods, but also the robustness w.r.t. click patterns.

1 papers0 benchmarksActions, Images, Interactive, Tables, Tabular

AV-HV interaction at Intersections (Interaction dataset of automated vehicles and human-driven vehicles at unsignalized intersections)

This is a processed dataset comprising the interactions of automated vehicles and human-driven vehicles at unsignalized intersections, extracted from the Waymo Open Motion Dataset and Lyft Level 5 Dataset.

1 papers0 benchmarks

SAT-MTB-VSR

SAT-MTB-VSR is a large-scale dataset for satellite video super-resolution made from original videos of Jilin-1, which is a subset of the satellite video multitasking dataset SAT-MTB. The dataset is cropped from 18 videos captured by the Jilin-1 video satellite, covering a wide range of terrains, such as cities, docks, airports, suburbs, forests, and deserts, with a resolution of about 1 m. And the videos contain dynamic scenes, such as moving cars, airplanes, trains, and ships, which test the ability of the VSR method to deal with moving targets of different sizes and speeds. At the same time, due to the motion of the satellite, the video contains changes in viewing angle and lighting.

1 papers11 benchmarksImages

difficult retrieval

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

ReALFRED

realfred is an embodied instruction following benchmark.

1 papers0 benchmarksImages, Texts

Driving Weather

A synthetic dataset including driving under adverse weather conditions | Autonomous Driving

1 papers0 benchmarksImages

HeightCeleb

Prediction of a speaker's height is of interest in fields such as voice forensics, surveillance, and automatic speaker profiling. HeightCeleb is an extension of Voxceleb that includes height information for all 1251 speakers. The height data was extracted automatically from publicly available sources. The purpose of this dataset is to enable the research community to leverage freely available speaker embedding extractors, pre-trained on VoxCeleb, to develop more accurate speaker height estimators.

1 papers0 benchmarksBiomedical

CAS-VSR-S101

A new large-scale, in-thewild Mandarin dataset, CAS-VSR-S101 with 101.1 hours of data. The videos are sourced from broadcast news and conversational programs in Chinese, covering a highly diverse set of topics, speakers and filming conditions. The lengths of the utterances are naturally distributed between 0.01s and 10.57s, and image qualities and resolutions vary. News accounts for 82.4% of the programs. 70.4% of the utterances depict news anchors, hosts and correspondents, while 29.6% are those of interviewees and guests. In addition, at a ratio of approximately 1.5 : 1, male and female appearances are relatively balanced. It is divided into train, validation and test sets by TV channels to minimize speaker overlap, and at a ratio of roughly 8 : 1 : 1.5 in terms of duration. The validation and test sets are composed of programs broadcast on provincial TV channels. The dataset is available for academic use under a license.

1 papers4 benchmarksAudio, Speech, Texts, Videos
PreviousPage 523 of 1000Next