19,997 machine learning datasets
19,997 dataset results
Analogical reasoning is fundamental to human cognition and holds an important place in various fields. However, previous studies mainly focus on single-modal analogical reasoning and ignore taking advantage of structure knowledge. We introduce the new task of multimodal analogical reasoning over knowledge graphs, which requires multimodal reasoning ability with the help of background knowledge. Our dataset MARS contains 10,685 training, 1,228 validation and 1,415 test instances.
LiDAR-CS is a dataset for 3D object detection in real traffic. It contains 84,000 point cloud frames under 6 groups of different sensors but with same corresponding scenarios, captured from hybrid realistic LivDAR simulator.
The data set is based on roughly 6,000 coffee bean review spublished on the website Coffereviews going back to 1997. All of these reviews are scored with the q-grading scale (Coffee Review, 2021)
Processed CMIP5 data used for testing the CNN-LSTM model. Details in Zenodo description
The 'Me 163' was a Second World War fighter airplane and a result of the German air force secret developments. One of these airplanes is currently owned and displayed in the historic aircraft exhibition of the 'Deutsches Museum' in Munich, Germany. To gain insights with respect to its history, design and state of preservation, a complete CT scan was obtained using an industrial XXL-computer tomography scanner at Fraunhofer EZRT .
Motion Capture Data for Hand Motion Embodiment contains demonstrations of different hand motion recorded with the Qualisys MOCAP system.
We release 3.691 earning call transcripts and also annotated data set, labeled particularly for the digital strategy maturity by linguists. https://github.com/hpataci/DigiCall
dataset (balanced) of 200 images consists of three classes - False Smut, Neck Blast and healthy grain class. Some of these images contain both diseases together. Field data collected under the supervision of staff from the Bangladesh Rice Research Institute (BRRI).
This data set encompasses 104 images and transcriptions of digital images of original charters from the Cistercian abbey Fontenay in Burgundy (France), dating mainly from the 12th c. and until 1213. The original data set was created as part of the ANR ORIFLAMMS (ANR-12-CORP-0010) project. Texts were transcribed in the original TEI-XML format, rendering both abbreviated and expanded forms of the original text. The alignment data was produced by merging coordinates created through the Oriflamms software. A new version was prepared in March-June 2022 as part of the research for the following paper: Camps, Jean-Baptiste, Chahan Vidal-Gorène, Dominique Stutzmann, Marguerite Vernet, and Ariane Pinche. « Data Diversity in Handwritten Text Recognition: Challenge or Opportunity? » In Digital Humanities 2022. Conference Abstracts (The University of Tokyo, Japan, 25-29 July 2022), published by DH2022 Local Organizing Committee, 160‑65. Tokyo, 2022.
The dataset consists of random electromagnetic scatterers and their associated fields when illuminated by a 1000nm plane-wave illumination. The scatterers have a refractive index of $n=1.5$ and are surrounded by air ($n=1.0$), and the side-length of the simulated area is 5.12 microns.
Dataset for User Verification part of MotionID: Human Authentication Approach. Data type: bin (should be converted by attached notebook). ~50 hours of IMU (Inertial Measurement Units) data for one specific motion pattern, provided by 101 users.
Dataset (part 1/3) for Motion Patterns Identification part of MotionID: Human Authentication Approach. Data type: bin (should be converted by attached notebook).
Dataset (part 2/3) for Motion Patterns Identification part of MotionID: Human Authentication Approach. Data type: bin (should be converted by attached notebook).
Dataset (part 3/3) for Motion Patterns Identification part of MotionID: Human Authentication Approach. Data type: bin (should be converted by attached notebook).
IISc VEED consists of 200 diverse indoor and outdoor scenes (see samples below). The videos are rendered with blender and the blend files are obtained for the scenes mainly from blendswap and turbosquid. 4 different camera trajectories are added to each scene and thus we have a total of 800 videos. The videos are rendered at full HD resolution (1920 x 1080) and at 30fps and contain 12 frames each.
The WUSTL-EHMS-2020 dataset was created using a real-time Enhanced Healthcare Monitoring System (EHMS) testbed [1]. This testbed collects both the network flow metrics and patients' biometrics due to the scarcity of a dataset that combines these biometrics.
We present the APRICOT-Mask dataset, which augments the APRICOT dataset with pixel-level annotations of adversarial patches. We hope APRICOT-Mask along with the APRICOT dataset can facilitate the research in building defenses against physical patch attacks, especially patch detection and removal techniques.
We manually annotate 800 sentences from 80 documents in two domains (Healthcare and Transportation) to form a DocOIE dataset for evaluation.
This dataset is used for MPLP considering time windows constraints of customers and parking space. To randomly generate the dataset, please visit the link: https://github.com/Yubin-Liu/Hybrid-Q-Learning-Network-Approach-for-MPLP.
Polysomnography (PSG) is a fundamental diagnostical method for the detection of Obstructive Sleep Apnea Syndrome (OSAS). Historically, trained physicians have been manually identifying OSAS episodes in individuals based on PSG recordings. Such a task is highly important for stroke patients, since in such cases OSAS is linked to higher mortality and worse neurological deficits. Unfortunately, the number of strokes per day vastly outnumbers the availability of polysomnographs and dedicated healthcare professionals. The data in this work pertains to 30 patients that were admitted to the stroke unit of the Udine University Hospital, Italy. Unlike previous studies, exclusion criteria are minimal. As a result, data are strongly affected by noise, and individuals may suffer from several comorbidities. Each patient instance is composed of overnight vital signs data deriving from multi-channel ECG, photoplethysmography and polysomnography, and related domain expert’s OSAS annotations. The datas