19,997 machine learning datasets
19,997 dataset results
Verified Smart Contracts is a dataset of real Ethereum smart contracts, containing both Solidity and Vyper source code. It consists of every deployed Ethereum smart contract as of 1st of April 2022, whose been verified on Etherscan and has a least one transaction. A total of 186,397 unique smart contracts are provided, filtered down from 2,217,692 smart contracts. The dataset contains 53,843,305 lines of code.
Verified Smart Contracts Code Comments is a dataset of real Ethereum smart contract functions, containing "code, comment" pairs of both Solidity and Vyper source code. The dataset is based on every deployed Ethereum smart contract as of 1st of April 2022, whose been verified on Etherscan and has a least one transaction. A total of 1,541,370 smart contract functions are provided, parsed from 186,397 unique smart contracts, filtered down from 2,217,692 smart contracts.
Vulnerable Verified Smart Contracts is a dataset of real vulnerable Ethereum smart contracts. Based on the manually labeled Benchmark dataset of Solidity smart contracts. A total of 609 vulnerable contracts are provided, containing 1,117 vulnerabilities.
Spatial TRAnsformation for virtual Try-on (STRAT) dataset contains three subdatasets: STRAT-glasses, STRAT-hat, and STRAT-tie, which correspond to "glasses try-on", "hat try-on", and "tie try-on" respectively. In each subdataset, the training set has 2000 pairs of foregrounds (accessories) and backgrounds (human faces or portrait images), while the test set has 1000 pairs of foregrounds and backgrounds. For each pairwise sample, both the vertice coordinates and warpping parameters of the foreground for each pairwise are provided for supervised learning and evaluation of spatial transformation.
This dataset comprises 16 499 images with 42 classes encompassing the most popular Central Asian cuisine consumed locally.
Kinematics Dataset for the NICOL robot (Neuro-inspired Collaborator). Data is intended for Training and Testing of inverse kinematics applications.
The dataset is specifically constructed for the library-oriented code generation task, which are constructed in the paper “CodeGen4Libs: A Two-Stage Approach for Library-Oriented Code Generation”.
This dataset encompasses 265 speeches (over 200,000 tokens) from the German Bundestag, primarily from the 19th legislative term (2017-2021), given by 195 distinct speakers representing 6 political parties.
Context The database contains wav recordings from the same optical sensor inserted in-turn into six insectary boxes containing only one mosquito species of both sexes (about 200-300 flying mosquitoes in each cage). As the mosquitoes fly randomly through the sensor their wingbeat partially occludes the light from the transmitter to the receiver. The light fluctuation recorded is modulated by the wingbeat of the insect. The resulting signal is pseudo-acoustic, meaning that it sounds exactly like a microphone recording but has been acquired using optical means (however, not vision based). Insect Biometrics, in the context of our work, is a measurable behavioral characteristic of flying insects. Biometric identifiers are related to the shape of the body (main body size, wing shape, wingbeat frequency, pattern movement of the wings). Biometric identification methods use biometric characteristics or traits to verify species/sex identities when insects access endpoint traps following a bait.
As part of our policy to openly share all data from this project, we have included a downloadable package comprising all acoustic data collected over the course of this work. This includes acoustic recordings from 20 different species of mosquitoes, using a variety of mobile phones for each. This data can be downloaded from the online repository on dryad.org. The supplementary audio files are not included in this package, and may be downloaded separately.
The dataset contains sentences from Amazon customer reviews (sampled from Amazon product review dataset) annotated for counterfactual detection (CFD) binary classification. Counterfactual statements describe events that did not or cannot take place. Counterfactual statements may be identified as statements of the form – If p was true, then q would be true (i.e. assertions whose antecedent (p) and consequent (q) are known or assumed to be false).
We use Kubric and ESIM simulator to make our EKubric dataset, which has 15,367 RGB-PointCloud-Event pairs with annotations (including optical flow, scene flow, surface normal, semantic segmentation and object coordinates ground truths).
This is a simulated dataset for force prediction
The instances were drawn randomly from a database of 7 outdoor images. The images were handsegmented to create a classification for every pixel. Each instance is a 3x3 region.
The task is to train a network to discriminate between sonar signals bounced off a metal cylinder and those bounced off a roughly cylindrical rock.
Robot@Home2, is an enhanced version aimed at improving usability and functionality for developing and testing mobile robotics and computer vision algorithms. Robot@Home2 consists of three main components. Firstly, a relational database that states the contextual information and data links, compatible with Standard Query Language. Secondly,a Python package for managing the database, including downloading, querying, and interfacing functions. Finally, learning resources in the form of Jupyter notebooks, runnable locally or on the Google Colab platform, enabling users to explore the dataset without local installations. These freely available tools are expected to enhance the ease of exploiting the Robot@Home dataset and accelerate research in computer vision and robotics.
Data for novelty and its impact detection in scientific publications from Microsoft Academic Graph (now OpenAlex)
InfraParis is a novel and versatile dataset supporting multiple tasks across three modalities: RGB, depth, and infrared. From the city to the suburbs, it contains a variety of styles in different areas of the greater Paris area, providing rich semantic information. InfraParis contains 7301 images with bounding boxes and full semantic (19 classes) annotations. We assess various state-of-the-art baseline techniques, encompassing models for the tasks of semantic segmentation, object detection, and depth estimation.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Two separate datasets of calibration runs in front of a calibration board: - 4IMUs+3Cams -4IMUs+4Cams