19,997 machine learning datasets
19,997 dataset results
We introduce an open-source physical-layer dataset of Bluetooth Low Energy (BLE) IoT sensor devices recorded in an anechoic chamber using USRP x310. With a 100Msps sampling rate, it covers the entire BLE spectrum, featuring on-body and off-body scenarios with 13 BLE devices (ESP32s) from the same manufacturer. the goal is to study the physical layer characteristics of both on-body and off-body signals. The dataset is also available through MongoDB with a Python tool for analysis; for more details, please visit our GitHub page. https://github.com/mkashani-phd/BLEWBAN_Dataset
This dataset called Indoor Lodz University of Technology Point Cloud Dataset (InLUT3D) is a point cloud set tailored for real object classification and both semantic and instance segmentation tasks. Comprising of 321 scans, some areas in the dataset are covered by multiple scans. All of them are captured using the Leica BLK360 scanner.
The Temporal Logic Video (TLV) Dataset addresses the scarcity of state-of-the-art video datasets for long-horizon, temporally extended activity and object detection. It comprises two main components:
The MS-EVS Dataset is the first large-scale event-based dataset for face detection.
The aim of KFall dataset is to contribute technology development for elderly fall detection and injury prevention. It was acquired by 32 young subjects with 21 types of activities of daily living (ADLs) and 15 types of falls from an inertial sensor attached on low back. In total, it contains 5075 motion files with 2729 ADL motions and 2346 fall motions. In addition, for each fall motion, ready-to-use fall labels (fall initialization and fall impact moment) based on synchronized video references were also included.
This dataset is comprised of the dynamic analysis reports generated by CAPEv2, from both malware and goodware. We source the goodware as they do in Dambra et al. (https://arxiv.org/abs/2307.14657), where trough the community-maintained packages of Chocolatey they create a dataset that spans 2012 to 2020. The malware are sourced from VirusTotal, namely samples of Portable Executable from 2017 - 2020 that they release for academic purposes. In total, the dataset we assembled contains 26,200 PE samples: 8,600 (33\%) goodware and 17,675 (67\%) malware.
The Lusitano dataset was collected over a 3-month period, spanning from January to March, from Paulo de Oliveira, S.A., a prominent textile company, based in Covilhã, Portugal, renowned for its innovative contributions to the textile industry. To collect the images for the dataset, we placed one camera in front of a fabric inspection machine, along with a strong and nearly uniform light source. This dataset comprises 4096 × 1024 images, captured by an industrial-grade Teledyne Dalsa Linea camera. The camera’s high resolution and precision ensure the accurate depiction of textile samples, with the level of detail necessary for defect analysis. None of the defects depicted in this dataset are artificially generated; they stem from genuine occurrences observed during this collection period, and thus represent real-world challenges encountered in textile production processes. The dataset also showcases normal images. We announce two folders, train and test in the same folder architecture,
We release a large-scale Multi-Scenario Multi-Modal CTR dataset named AntM2C, built from real industrial data from Alipay. This dataset offers an impressive breadth and depth of information, covering CTR data from four diverse business scenarios, including advertisements, consumer coupons, mini-programs, and videos. Unlike existing datasets, AntM2C provides not only ID-based features but also five textual features and one image feature for both users and items, supporting more delicate multi-modal CTR prediction.
This dataset contains over 12K+ questions and symptoms related to various common diseases in Vietnamese. It's designed to aid in the classification of medical symptoms and provide preliminary disease identification. The dataset covers a wide range of diseases, including cardiovascular, digestive, neurological, dermatological, endocrine, and others.
we construct the ViMirr dataset, which has 19,255 frames from 276 videos
Perception systems of autonomous vehicles are susceptible to occlusion, especially when examined from a vehicle-centric perspective. Such occlusion can lead to overlooked object detections, e.g., larger vehicles such as trucks or buses may create blind spots where cyclists or pedestrians could be obscured, accentuating the safety concerns associated with such perception system limitations. To mitigate these challenges, the vehicle-to-everything (V2X) paradigm suggests employing an infrastructure-side perception system (IPS) to complement autonomous vehicles with a broader perceptual scope. Nevertheless, the scarcity of real-world 3D infrastructure-side datasets constrains the advancement of V2X technologies. To bridge these gaps, this paper introduces a new 3D infrastructure-side collaborative perception dataset, abbreviated as inscope. Notably, InScope is the first dataset dedicated to addressing occlusion challenges by strategically deploying multiple-position Light Detection and Ran
Source: MLHME-38K
Demo dataset with 5 partial and complete 3D shapes of potato tubers.
sheeep ruminate behavior image and label
sheeep ruminate behavior image and label
The Internet of Things (IoT) is omnipresent, exposing a large number of devices that often lack security controls to the public Internet. In the modern world, many everyday processes depend on these devices, and their service outage could lead to catastrophic consequences. There are many Deep Packet Inspection (DPI) based intrusion detection systems (IDS). However, their linear computational complexity induced by the event-driven nature poses a power-demanding obstacle in resource-constrained IoT environments. In this paper, we shift away from the traditional IDS as we introduce a novel and lightweight framework, relying on a time-driven algorithm to detect Distributed Denial of Service (DDoS) attacks by employing Machine Learning (ML) algorithms leveraging the newly engineered features containing system and network utilization information. These features are periodically generated, and there are only ten of them, resulting in a low and constant algorithmic complexity. Moreover, we lev
The Internet of Things (IoT) is omnipresent, exposing a large number of devices that often lack security controls to the public Internet. In the modern world, many everyday processes depend on these devices, and their service outage could lead to catastrophic consequences. There are many Deep Packet Inspection (DPI) based intrusion detection systems (IDS). However, their linear computational complexity induced by the event-driven nature poses a power-demanding obstacle in resource-constrained IoT environments. In this paper, we shift away from the traditional IDS as we introduce a novel and lightweight framework, relying on a time-driven algorithm to detect Distributed Denial of Service (DDoS) attacks by employing Machine Learning (ML) algorithms leveraging the newly engineered features containing system and network utilization information. These features are periodically generated, and there are only ten of them, resulting in a low and constant algorithmic complexity. Moreover, we lev
The Internet of Things (IoT) is omnipresent, exposing a large number of devices that often lack security controls to the public Internet. In the modern world, many everyday processes depend on these devices, and their service outage could lead to catastrophic consequences. There are many Deep Packet Inspection (DPI) based intrusion detection systems (IDS). However, their linear computational complexity induced by the event-driven nature poses a power-demanding obstacle in resource-constrained IoT environments. In this paper, we shift away from the traditional IDS as we introduce a novel and lightweight framework, relying on a time-driven algorithm to detect Distributed Denial of Service (DDoS) attacks by employing Machine Learning (ML) algorithms leveraging the newly engineered features containing system and network utilization information. These features are periodically generated, and there are only ten of them, resulting in a low and constant algorithmic complexity. Moreover, we lev
The CHAINS dataset is utilized to evaluate the proposed system. This dataset includes both normal speech, referred to as solo, and whispered speech, known as whsp. The dataset comprises recordings from 36 speakers, dominantly with Irish dialects. To develop a model capable of accurately capturing the acoustic features of the Irish dialect, the normal speech portion of this dataset is used for system training, while the whispered speech portion is employed for evaluation.
The deliberate manipulation of public opinion, especially through altered images, poses a significant danger to society. To fight this issue on a technical level we support the research community by releasing the Digital Forensics 2023 (DF2023) training and validation dataset.