19,997 machine learning datasets
19,997 dataset results
Contains correlation data for 119,384 column pairs, taken from 3,952 data sets, including Pearson correlation, Spearman correlation, and Theil's U. This data can be used, e.g., for approaches that predict column correlation based on column properties, including column names.
We construct a large-scale Heterogeneous Graph benchmark dataset named UniKG from Wikidata. UniKG contains 77.31 million multi-attribute entities labels by 2000 classes, 564 million directed edges annotated by 2082 diverse association types, which significantly surpasses the scale of existing homogeneous graph datasets. UniKG have the capability to facilitate downstream task of diverse domains.
The CATH (Class, Architecture, Topology, Homology) [65] database is a comprehensive resource for protein structure classification that hierarchical group proteins based on their structural features. The database defines classes based on topological similarities, architectures based on the arrangement of secondary structure elements, topologies based on the connectivity of secondary structure elements, and homologous domains based on sequence similarity.
The CATH (Class, Architecture, Topology, Homology) [65] database is a comprehensive resource for protein structure classification that hierarchical group proteins based on their structural features. The database defines classes based on topological similarities, architectures based on the arrangement of secondary structure elements, topologies based on the connectivity of secondary structure elements, and homologous domains based on sequence similarity. This results in a training set of 16,153 structures, a validation set of 1,457 structures, and a test set of 1,797 structures. Note that the curated CATH dataset contains only single-chain structures and does not consider the case of designing multi-chain proteins.
Repository of containerized services that can be migrated through UMS. The dataset includes containers of the UMS platform plus sample containerized services that can be live migrated using UMS. Specifically, the dataset includes containers for the following two services: 1. Memhog application: this is a containerized service to check the impact of memory footprint of the containers upon live migration. 2. Yolo v3-Tiny application: this is a containerized service with a real-world application for object detection. This will help users to examine UMS under real-world settings.
57 stock videos from Pexels, predominantly covering road scenes which involve minimal distortion.
CSV file with a list of all examined OWL reasoners. For each item, information on usability and maintenance status, project pages, source code repositories and related documentation was gathered.
FROG is a 2D LiDAR dataset with annotations for people detectors. It consists of 6 fully annotated sequences, and 30 total hours of recordings at the Royal Alcázar of Seville (Spain). The main motivation of this dataset is providing higher quality data with a richer variety of crowded scenarios, for improving the field of people detectors based on knee-high 2D range finders.
All data is from one continuous EEG measurement with the Emotiv EEG Neuroheadset. The duration of the measurement was 117 seconds. The eye state was detected via a camera during the EEG measurement and added later manually to the file after analysing the video frames. '1' indicates the eye-closed and '0' the eye-open state. All values are in chronological order with the first measured value at the top of the data.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
The Failure Mode Classification dataset released in the paper "MWO2KG and Echidna: Constructing and exploring knowledge graphs from maintenance data" by Stewart et al. The goal is to label a given observation (made by a maintainer) with the corresponding Failure Mode Code.
Continuous EEG activity was recorded from each member of the dyad using an ActiveTwo head cap and the ActiveTwo Biosemi system (BioSemi, Amsterdam, Netherlands). Recordings were collected from 64 Ag-AgCl scalp electrodes and from bilateral mastoids. Two electrodes were placed next to each other 1 cm below the right eye to record eye-blink responses. A ground electrode was established by BioSemi’s common Mode Sense active electrode and Driven Right Leg passive electrode. EEG activity was digitized with ActiView software (BioSemi) and sampled at 2048 Hz. Data was downsampled post-acquisition and analyzed at 512 Hz.
Synthetic Speech Attribution Dataset.
This synthetic event dataset is used in Robust e-NeRF to study the collective effect of camera speed profile, contrast threshold variation and refractory period on the quality of NeRF reconstruction from a moving event camera. It is simulated using an improved version of ESIM with three different camera configurations of increasing difficulty levels (i.e. easy, medium and hard) on seven Realistic Synthetic $360^{\circ}$ scenes (adopted in the synthetic experiments of NeRF), resulting in a total of 21 sequence recordings. Please refer to the Robust e-NeRF paper for more details.
In this dataset an uppertorso humanoid robot with 7-DOF arm explored 100 different objects belonging to 20 different categories using 10 behaviors: Look, Crush, Grasp, Hold, Lift, Drop, Poke, Push, Shake and Tap.
The data can be found in the Data folder, which contains two files:
This dataset contains the ground truth for urban changes occurred in Mariupol, Ukraine for the time frame 2017-2020. This is useful for transferring the urban change monitoring network ERCNN-DRS (https://github.com/It4innovations/ERCNN-DRS_urban_change_monitoring) to that region.
The CapMIT1003 database contains captions and clicks collected for images from the MIT1003 database, for which reference eye scanpath are available. The database is distributed as a single SQLite3 database named capmit1003.db. For convenience, a lightweight Python class to access the database is provided in the official repository
we introduce a large-scale and diverse symbolic melody dataset called MelodyNet that contains more than 0.4 million melody pieces extracted from approximately 1.6 million songs. MelodyNet is used for large-scale pre-training and domain-specific n-gram lexicon construction.
We present a new collection of 1,981 Vega-Lite specifications, which is used to demonstrate the generalizability and viability of our NL generation framework. This collection is the largest set of human-generated charts obtained from GitHub to date. It covers varying levels of complexity from a simple line chart without any interaction to a chart with four plots where data points are linked with selection interactions. Compared to the benchmarks, our dataset shows the highest average pairwise edit distance between specifications, which proves that the charts are highly diverse from one another. Moreover, it contains the largest number of charts with composite views, interactions (e.g., tooltips, panning & zooming, and linking), and diverse chart types (e.g., map, grid & matrix, diagram, etc.).