19,997 machine learning datasets
19,997 dataset results
RepoIMU T-stick The RepoIMU T-stick is a small, low-cost, and high-performance inertial measurement unit (IMU) that can be used for a wide range of applications. The RepoIMU T-stick is a 9-axis IMU that measures the acceleration, angular velocity, and magnetic field. This database contains two separate sets of experiments recorded with a T-stick and a pendulum. A total of 29 trials were collected on the T-stick, and each trial lasted approximately 90 seconds. As the name suggests, the IMU is attached to a T-shaped stick equipped with six reflective markers. Each experiment consists of slow or fast rotation around a principal sensor axis or translation along a principal sensor axis. In this scenario, the data from the Vicon Nexus OMC system and the XSens MTi IMU are synchronized and provided at a frequency of 100 Hz. The authors clearly state that the IMU coordinate system and the ground trace are not aligned and propose a method to compensate for one of the two required rotations base
ISOD contains 2,000 manually labelled RGB-D images from 20 diverse sites, each featuring over 30 types of small objects randomly placed amidst the items already present in the scenes. These objects, typically ≤3cm in height, include LEGO blocks, rags, slippers, gloves, shoes, cables, crayons, chalk, glasses, smartphones (and their cases), fake banana peels, fake pet waste, and piles of toilet paper, among others. These items were chosen because they either threaten the safe operation of indoor mobile robots or create messes if run over.
UniRef90 is generated by clustering UniRef100 seed sequences.
The Belfort dataset This dataset includes minutes of Belfort municipal council drawn up between 1790 and 1946. Documents include deliberations, lists of councillors, convocations, and agendas. It includes 24,105 text-line images that were automatically detected from pages. Up to 4 transcriptions are available for each line image: two from humans, and two from automatic models.
OCTScenes contains 5000 tabletop scenes with a total of 15 everyday objects. Each scene is captured in 60 frames covering a 360-degree perspective.
MSVD-Indonesian is derived from the MSVD dataset, which is obtained with the help of a machine translation service. This dataset can be used for multimodal video-text tasks, including text-to-video retrieval, video-to-text retrieval, and video captioning. Same as the original English dataset, the MSVD-Indonesian dataset contains about 80k video-text pairs.
This dataset contains the pre-generated dataset referenced in the GenPlot Paper.
The data contains CSV files with anonymized user names, tweet texts, vaccine stance, cumulative score for the vaccine stance, location, and topic information. The file named all_predicted_cumulative_stance.csv contains all the tweets, scores, and classifications. We have broken this file into two separate files named demotivate_cumulative_stance.csv and motivate_cumulative_stance.csv, containing the demotivating and motivating tweets, respectively. We used these two files in the visualization tool presented at: https://ashiqur-rony.github.io/visualize-covid-stance/
The whole UCF-Crime dataset consists of real-world 240 × 320 RGB videos with 13 realistic anomaly types such as explosion, road accident, burglary, etc., and normal examples. The CPD specific requires a change in data distribution. We suppose that explosions and road accidents correspond to such a scenario, while most other types correspond to point anomalies. For example, data, obviously, com from a normal regime before the explosion. After it, we can see fire and smoke, which last for some time. Thus, the first moment when an explosion appears is a change point. Along with a volunteer, the authors carefully labelled chosen anomaly types. Their opinions were averaged. We provide the obtained markup, so other researchers can use it to validate their CPD algorithm for video.
Annotation corpus of cybersecurity event in news articles The corpus contains 1000 annotation and source files. Our cybersecurity focused on five event types: Databreach, Phishing, Ransom, Discover, and Patch.
This is a dataset of regular expressions collected from regex101.com. It is not made directly available, but can be crawled from regex101.
Multi-Person Interaction Motion (MI-Motion) Dataset includes skeleton sequences of multiple individuals collected by motion capture systems and refined and synthesized using a game engine. The dataset contains 167k frames of interacting people's skeleton poses and is categorized into 5 different activity scenes.
DeepGraviLens is a data set of simulated gravitational lenses consisting of images associated with brightness variation time series. In this dataset, both non-transient and transient phenomena (supernovae explosions) are simulated.
Overview This is a dataset of blood cells photos.
L3Cube-MahaCorpus is a Marathi monolingual data set scraped from different internet sources. We expand the existing Marathi monolingual corpus with 24.8M sentences and 289M tokens. We also present, MahaBERT, MahaAlBERT, and MahaRoBerta all BERT-based masked language models, and MahaFT, the fast text word embeddings both trained on full Marathi corpus with 752M tokens.
PTVD is a plot-oriented multimodal dataset in the TV domain. It is also the first non-English dataset of its kind. Additionally, PTVD contains more than 26 million bullet screen comments (BSCs), powering large-scale pre-training.
EgoISM-HOI is a new multimodal dataset composed of synthetic and real images of egocentric human-objects interactions in an industrial environment with rich annotations of hands and objects. EgoISM-HOI contains a total of 39,304 RGB images, 23,356 depth maps and instance segmentation masks, 59,860 hand annotations, 237,985 object instances across 19 object categories and 35,416 egocentric human-object interactions.
Generated for further pre-training pre-trained models like BERT, RoBERTa, ALBERT, DeBERTa, etc.. in order to get stronger logical reasoning ability.
WDC Block is a benchmark for comparing the performance of blocking methods that are used as part of entity resolution pipelines.
3D-Speaker is a large-scale speech corpus designed to facilitate the research of speech representation disentanglement. 3DSpeaker contains over 10,000 speakers, each of whom are simultaneously recorded by multiple Devices, locating at different Distances, and some speakers are speaking multiple Dialects. The controlled combinations of multi-dimensional audio data yield a matrix of a diverse blend of speech representations entanglement, thereby motivating intriguing methods to untangle them.