19,997 machine learning datasets
19,997 dataset results
The scene derives from photo-realistic HM3D datasets. Our dataset offers a wide variety of environments especially for Social Navigation tasks, with carefully calibrated human density, incorporating realistic human motions and natural movement patterns. These features ensure balanced interaction dynamics across diverse scenes, facilitating the development of more effective social navigation algorithms.
The scene derives from photo-realistic MP3D datasets. Our dataset offers a wide variety of environments especially for Social Navigation tasks, with carefully calibrated human density, incorporating realistic human motions and natural movement patterns. These features ensure balanced interaction dynamics across diverse scenes, facilitating the development of more effective social navigation algorithms.
The NuiSI dataset contains skeleton tracking trajectories of Human Interaction Partners performing a variety of physically interactive behaviors (waving, handshaking, rocket fistbump, parachute fistbump) with each other. This is inspired by the dataset in Bütepage et al. "Imitating by generating: Deep generative models for imitation of interactive tasks." Frontiers in Robotics and AI (2020) wherein they capture a dataset with rokoko motion capture suits. Instead we track the skeletons of the interaction partner with Intel Realsense cameras using Nuitrack, for a more realistic scenario, with noise coming from the depth sensor, the skeleton tracking and some partial occlusions. This makes it more representative of real world interactions with a Robot equipped with an RGBD camera. T This dataset is used in our papers for training Interaction models for Human-Robot Interaction with a humanoid social robot. If you find the dataset useful in your work, please cite our paper:
Dataset used in Bütepage, Judith, et al. "Imitating by generating: Deep generative models for imitation of interactive tasks." Frontiers in Robotics and AI 7 (2020): 47.
Overnight is a dataset for semantic parsing in eight domains.
This dataset is a collection of marxist fragments mixed and cut randomly from the Marxist archive (marxists.org).
The dataset SCARED-C is introduced in the context of assessing robustness in endoscopic depth prediction models. It is part of the EndoDepth benchmark, which is designed to evaluate the performance of monocular depth prediction models specifically for endoscopic scenarios. The dataset features 16 different types of image corruptions, each with five levels of severity, encompassing challenges like lens distortion, resolution alterations, specular reflection, and color changes that are typical in endoscopic imaging. The ground truth is on the original testing set of SCARED.
Vision-language supervised fine-tuning effectively enhances VLLM performance, but existing visual instruction tuning datasets have limitations:
A collection of various NLP datasets in Assamese. These datasets are split into two: pre-training corpora and fine-tuning datasets.
We introduce a video dataset Bukva for Russian Dactyl Recognition task. Bukva dataset size is about 27 GB, and it contains 3757 RGB videos with more than 101 samples for each RSL alphabet sign, including dynamic ones. The dataset is divided into training set and test set by subject user_id. The training set includes 3097 videos, and the test set includes 660 videos. The total video recording time is ~4 hours. About 17% of the videos are recorded in HD format, and 70% of the videos are in FullHD resolution.
A Racial Fairness Benchmark Dataset for Face Forgery Detection.
JamPatoisNLI provides the first dataset for natural language inference in a creole language, Jamaican Patois. Many of the most-spoken low-resource languages are creoles. These languages commonly have a lexicon derived from a major world language and a distinctive grammar reflecting the languages of the original speakers and the process of language birth by creolization. This gives them a distinctive place in exploring the effectiveness of transfer from large monolingual or multilingual pretrained models. While our work, along with previous work, shows that transfer from these models to low-resource languages that are unrelated to languages in their training set is not very effective, we would expect stronger results from transfer to creoles. Indeed, our experiments show considerably better results from few-shot learning of JamPatoisNLI than for such unrelated languages, and help us begin to understand how the unique relationship between creoles and their high-resource base languages af
We conducted a large crowdsourcing study of click patterns in an interactive segmentation scenario and collected 475K real-user clicks. Drawing on ideas from saliency tasks, we develop a clickability model that enables sampling clicks, which closely resemble actual user inputs. Using our model and dataset, we propose RClicks benchmark for a comprehensive comparison of existing interactive segmentation methods on realistic clicks. Specifically, we evaluate not only the average quality of methods, but also the robustness w.r.t. click patterns.
This is a processed dataset comprising the interactions of automated vehicles and human-driven vehicles at unsignalized intersections, extracted from the Waymo Open Motion Dataset and Lyft Level 5 Dataset.
SAT-MTB-VSR is a large-scale dataset for satellite video super-resolution made from original videos of Jilin-1, which is a subset of the satellite video multitasking dataset SAT-MTB. The dataset is cropped from 18 videos captured by the Jilin-1 video satellite, covering a wide range of terrains, such as cities, docks, airports, suburbs, forests, and deserts, with a resolution of about 1 m. And the videos contain dynamic scenes, such as moving cars, airplanes, trains, and ships, which test the ability of the VSR method to deal with moving targets of different sizes and speeds. At the same time, due to the motion of the satellite, the video contains changes in viewing angle and lighting.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
realfred is an embodied instruction following benchmark.
A synthetic dataset including driving under adverse weather conditions | Autonomous Driving
Prediction of a speaker's height is of interest in fields such as voice forensics, surveillance, and automatic speaker profiling. HeightCeleb is an extension of Voxceleb that includes height information for all 1251 speakers. The height data was extracted automatically from publicly available sources. The purpose of this dataset is to enable the research community to leverage freely available speaker embedding extractors, pre-trained on VoxCeleb, to develop more accurate speaker height estimators.
A new large-scale, in-thewild Mandarin dataset, CAS-VSR-S101 with 101.1 hours of data. The videos are sourced from broadcast news and conversational programs in Chinese, covering a highly diverse set of topics, speakers and filming conditions. The lengths of the utterances are naturally distributed between 0.01s and 10.57s, and image qualities and resolutions vary. News accounts for 82.4% of the programs. 70.4% of the utterances depict news anchors, hosts and correspondents, while 29.6% are those of interviewees and guests. In addition, at a ratio of approximately 1.5 : 1, male and female appearances are relatively balanced. It is divided into train, validation and test sets by TV channels to minimize speaker overlap, and at a ratio of roughly 8 : 1 : 1.5 in terms of duration. The validation and test sets are composed of programs broadcast on provincial TV channels. The dataset is available for academic use under a license.