19,997 machine learning datasets
19,997 dataset results
This is the data regarding the pre-generated demonstration and experience replay for the proposed Deep-GRAIL algorithm. You are welcomed to generate your own replays based on your problems at hand.
The dataset contains procedurally generated images of transparent vessels containing liquid and objects . The data for each image includes segmentation maps, 3d depth maps, and normal maps of of the liquid or object inside the transparent vessel, and the vessel. In addition, the properties of the materials inside the containers are given(color/transparency/roughness/metalness). In addition, a natural image benchmark for the 3d/depth estimation of objects inside transparent containers is supplied. 3d models of the objects (GTLF) are also supplied.
PAGE contains 98,525 games played by 2,007 professional players and spans over 70 years. The dataset includes rich AI analysis results for each move.
Jericho Environment Commonsense Comprehension (JECC) is a dataset for commonsense reasoning. It consists of 29 games in multiple domains from the Jericho Environment hausknecht2019interactive.
NLI4Wills Corpus can be used to train transformers and sentence-transformer models for the validity evaluation of the legal will statements. Our dataset consists of ID numbers, three types of inputs (legal will statements, laws, and conditions) and classifications (support, refute, or unrelated).
Landsat 8 Collection 1 Tier 1 and Real-Time data DN values, representing scaled, calibrated at-sensor radiance.
Raw-Microscopy:
TempWikiBio is a new data-to-text generation dataset containing more than 4 millions of chronologically ordered revisions of biographical articles from English Wikipedia, each paired with structured personal profiles.
A dataset of real-world underwater videos annotated with multi-object tracking labels. The data was collected of the coast of the Big Island of Hawaii and the primary goal is to help scientists studying fish behavior, with the goal of conserving rare and beautiful fish species.
RealHDRTV dataset is the first real-world paired SDRTV-HDRTV dataset, which includes SDRTV-HDRTV pairs with 8K resolutions captured by a smartphone camera with the “SDR” and “HDR10” modes. To avoid possible misalignment, a professional steady tripod is used and only captured indoor or in controlled static scenes. After the acquisition, regions are cut out with obvious motions (10+ pixels) and light condition changes, and are cropped into 4K image pairs and a global 2D translation is used to align the cropped image pairs. Then, the pairs are removed which are still with obvious misalignment and get final 4K SDRTV-HDRTV pairs with misalignment no more than 1 pixel as labeled inference dataset.
slopt_fuzzbench_and_bandit_plot_data.tar.gz contains all plot_data of fuzzer instances that were run in the FuzzBench benchmark (Section 5.3) and Bandit Algorithm Comparison (Section 4.3).
slopt_magma_jsons.tar.gz contains the summary of the Magma benchmark (Section 5.4) as JSON files, which was generated by exp2json.py.
This data set contains over 600GB of multimodal data from a Mars analog mission, including accurate 6DoF outdoor ground truth, indoor-outdoor transitions with continuous cross-domain ground truth, and indoor data with Optitrack measurements as ground truth. With 26 flights and a combined distance of 2.5km, this data set provides you with various distinct challenges for testing and proofing your algorithms. The UAV carries 18 sensors, including a high-resolution navigation camera and a stereo camera with an overlapping field of view, two RTK GNSS sensors with centimeter accuracy, as well as three IMUs, placed at strategic locations: Hardware dampened at the center, off-center with a lever arm, and a 1kHz IMU rigidly attached to the UAV (in case you want to work with unfiltered data). The sensors are fully pre-calibrated, and the data set is ready to use. However, if you want to use your own calibration algorithms, then the raw calibration data is also ready for download. The cross-domai
Halpe-FullBody is a full body keypoints dataset where each person has annotated 136 keypoints, including 20 for body, 6 for feet, 42 for hands and 68 for face. It is designed for the task of whole body human pose estimation.
EventEA is an event-centric entity alignment dataset, harvested from EventKG, DBpedia and Wikidata.
NJH is a dataset of over 40,000 tweets about immigration from the US and UK, annotated with six labels for different aspects of incivility and intolerance. It is a more fine-grained multi-label approach to predicting incivility and hateful or intolerant content.
A synthetic sound mixture specification dataset for the Target Sound Extraction (TSE) task. Dataset samples consist of a .jams file specifying the mixture components, and a metadata file with target labels. Mixtures are 6 seconds long and contain 3-5 unique foreground sounds over a 6 second long background sound. Each sample is provided with 3 target labels, and sounds corresponding to all target labels are guaranteed to be present in the mixture. FSDKaggle2018 is used as the source for foreground sounds and TAU Urban Acoustic Scenes 2019 is used as the source for background sounds.
Visual Commonsense Immorality benchmark is a benchmark designed to evaluate commonsense immorality. It contains 2,172 immoral images for general and extensive immoral image detection.
AtyPict is a dataset of atypical sketch content designed for atypical sketch content detection tasks.
Insects are the most important global pollinator of crops and play a key role in maintaining the sustainability of natural ecosystems. Insect pollination monitoring and management are therefore essential for improving crop production and food security. Computer vision-facilitated pollinator monitoring can intensify data collection over what is feasible using manual approaches. We introduce a novel system to facilitate markerless data capture for insect counting, insect motion tracking, behaviour analysis and pollination prediction across large agricultural areas. Our system is comprised of edge computing multi-point video recording, offline automated multi-species insect counting, tracking and behavioural analysis. We implement and test our system on a commercial berry farm to demonstrate its capabilities.