19,997 machine learning datasets
19,997 dataset results
Description The molecules were randomly drawn from the largest set of isomers in the QM9 dataset [1] which consists of molecules with a fixed composition of atoms (C7O2H10) arranged in different chemically valid structures. It is an extension of the ismoer MD data used in [2].
Data used in paper entitled "Multiple Testing and Variable Selection along Least Angle Regression's path".
HYPE Dataset - Version 1.0.0
data/images:
DIT4BEARs Internship Project (at UiT-The Arctic University of Norway) Dataset
This is a challenging dataset from real courtrooms to predict the legal judgment in a reasonably encyclopedic manner by leveraging the genuine input of the case -- plaintiff's claims and court debate data, from which the case's facts are automatically recognized by comprehensively understanding the multi-role dialogues of the court debate, and then learnt to discriminate the claims so as to reach the final judgment through multi-task learning.
Synthetic graph classification datasets with the task of recognizing the connectivity of same-colored nodes in 4 graphs of varying topology.
The dataset contains entities from IMDB, TheMovieDB and TheTVDB with goldstandard matches between the sources. Due to the licensing of IMDB we provide a script to build the IMDB part of the dataset yourself.
This dataset contains a total of 11 variables. These are: 1. vectorprice: The value in local currency of the product 2. Exchange: The official exchange rate between USD and the local currency when data was extracted. 3. Usprice: The price of the product in USD 4. vectorsold: The number of items sold by the vendor when data was extracted. 5. vectorproduct: The name of the product sold by the vendor 6. country: The name of the country where the product was sold. 7. vectorquestions: The number of questions that the vendor received when data was extracted 8. goodfeedback: the number of positive feedbacks that the vendor received when data was extracted 9. neutralfeedback: the number of neutral feedback (neither positive nor negative) 10. badfeedback: the number of negative feedback. 11. Trust: Just the ratio between goodfeedback divided by goodfeedback + neutralfeedback + badfeedback
The Varied Emotion in Syntactically Uniform Speech (VESUS) repository is a lexically controlled database collected by the NSA lab. Here, actors read a semantically neutral script of words, phrases, and sentences with different emotional inflections. VESUS contains 252 distinct phrases, each read by 10 actors in 5 emotional states (neutral, angry, happy, sad, fearful).
This repository contains all the ensemble datasets (along with their meta-data) used in the manuscript "Wasserstein Distances, Geodesics and Barycenters of Merge Trees".
We introduce the CADNET dataset, which is an annotated collection of 3,317 3D Engineering models over 43 categories. Owing to the availability of large annotated datasets and also enough computational power in the form of GPUs, many deep learning-based solutions for object classification have been proposed of late, especially in the domain of images and graphical models. Nevertheless, very few solutions have been proposed for the task of functional classification of CAD models. Hence, for this research, CAD models have been collected from Engineering Shape Benchmark (ESB), National Design Repository (NDR), and augmented with newer models created using a modeling software to form a dataset - ‘CADNET’.
The dataset contains more than 100k code patch pairs extracted from open source projects on GitHub. Each pair comes with the erroneous and the fixed version of the corresponding code snippet. Instead of the whole file, the code snippets are extracted to focus on the problematic region (error line + other lines around it). For each sample, the repository name, the commit id, and the file names are provided so that one can access the complete files in case of interest.
OG RGB+D is a new gait recognition database called OG RGB+D database, which breaks through the limitation of other gait databases and includes multimodal gait data of various occlusions (self-occlusion, active occlusion, and passive occlusion) by a multiple synchronous Azure Kinect DK sensors data acquisition system (multi-Kinect SDAS) that can be also applied in security situations. Because Azure Kinect DK can simultaneously collect multimodal data to support different types of gait recognition algorithms, especially enables to effectively obtain camera-centric multi-person 3D poses, and multi view is better to deal with occlusion than single-view. In particular, the OG RGB+D database provides accurate silhouettes and the optimized human 3D joints data (OJ) by fusing data collected by multi-Kinects which are more accurate in human pose representation under occlusion.
CARLE is a life-like cellular automata simulator and reinforcement learning environment. CARLE is flexible, capable of simulating any of the 262,144 different rules defining Life-like cellular automaton universes. CARLE is also fast and can simulate automata universes at a rate of tens of thousands of steps per second through a combination of vectorization and GPU acceleration. Finally, CARLE is simple. Compared to high-fidelity physics simulators and video games designed for human players, CARLE's two-dimensional grid world offers a discrete, deterministic, and atomic universal playground, despite its complexity.
This data contains the election polls for the 2004, 2008, 2012, and 2016 US presidential election by state including data on undecided voter proportions.
DeCost, Hecht, Francis, Webler, Picard, and Holm. UHCSDB (Ultrahigh Carbon Steel micrograph DataBase): tools for exploring large heterogeneous microstructure datasets. accepted for publication in IMMI 2017 doi: 10.1007/s40192-017-0097-0
We present a data set from a first-principles study of amino-methylated and acetylated (capped) dipeptides of the 20 proteinogenic amino acids – including alternative possible side chain protonation states and their interactions with selected divalent cations (Ca$^{2+}$, Mg$^{2+}$ and Ba$^{2+}$. The data covers 21,909 stationary points on the respective potential-energy surfaces in a wide relative energy range of up to 4 eV (390 kJ/mol). Relevant properties of interest, like partial charges, were derived for the conformers.
This onotology is populated with the data from AADB2021 (https://dx.doi.org/10.17172/NOMAD/2021.02.10-1). Details can be found in the related article on arXiv.org: https://arxiv.org/abs/2107.08855
Dataset Files The official dataset files are hosted at https://dx.doi.org/10.21227/1y7r-am78.