TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Moh (SeyedMohammad Kashani)

We introduce an open-source physical-layer dataset of Bluetooth Low Energy (BLE) IoT sensor devices recorded in an anechoic chamber using USRP x310. With a 100Msps sampling rate, it covers the entire BLE spectrum, featuring on-body and off-body scenarios with 13 BLE devices (ESP32s) from the same manufacturer. the goal is to study the physical layer characteristics of both on-body and off-body signals. The dataset is also available through MongoDB with a Python tool for analysis; for more details, please visit our GitHub page. https://github.com/mkashani-phd/BLEWBAN_Dataset

0 papers0 benchmarksBiomedical, Time series

InLUT3D (Indoor Lodz University of Technology Point Cloud Dataset)

This dataset called Indoor Lodz University of Technology Point Cloud Dataset (InLUT3D) is a point cloud set tailored for real object classification and both semantic and instance segmentation tasks. Comprising of 321 scans, some areas in the dataset are covered by multiple scans. All of them are captured using the Leica BLK360 scanner.

0 papers0 benchmarks3D, Graphs, LiDAR, Point cloud

Temporal Logic Video (TLV) Dataset

The Temporal Logic Video (TLV) Dataset addresses the scarcity of state-of-the-art video datasets for long-horizon, temporally extended activity and object detection. It comprises two main components:

0 papers0 benchmarksImages, Videos

MS-EVS Dataset (Multispectral Event-based Face detection dataset)

The MS-EVS Dataset is the first large-scale event-based dataset for face detection.

0 papers0 benchmarksHyperspectral images, Images, Videos

KFall Dataset

The aim of KFall dataset is to contribute technology development for elderly fall detection and injury prevention. It was acquired by 32 young subjects with 21 types of activities of daily living (ADLs) and 15 types of falls from an inertial sensor attached on low back. In total, it contains 5075 motion files with 2729 ADL motions and 2346 fall motions. In addition, for each fall motion, ready-to-use fall labels (fall initialization and fall impact moment) based on synchronized video references were also included.

0 papers0 benchmarks

AutoRobust

This dataset is comprised of the dynamic analysis reports generated by CAPEv2, from both malware and goodware. We source the goodware as they do in Dambra et al. (https://arxiv.org/abs/2307.14657), where trough the community-maintained packages of Chocolatey they create a dataset that spans 2012 to 2020. The malware are sourced from VirusTotal, namely samples of Portable Executable from 2017 - 2020 that they release for academic purposes. In total, the dataset we assembled contains 26,200 PE samples: 8,600 (33\%) goodware and 17,675 (67\%) malware.

0 papers0 benchmarksTexts

Lusitano Fabric Defect Detection Dataset

The Lusitano dataset was collected over a 3-month period, spanning from January to March, from Paulo de Oliveira, S.A., a prominent textile company, based in Covilhã, Portugal, renowned for its innovative contributions to the textile industry. To collect the images for the dataset, we placed one camera in front of a fabric inspection machine, along with a strong and nearly uniform light source. This dataset comprises 4096 × 1024 images, captured by an industrial-grade Teledyne Dalsa Linea camera. The camera’s high resolution and precision ensure the accurate depiction of textile samples, with the level of detail necessary for defect analysis. None of the defects depicted in this dataset are artificially generated; they stem from genuine occurrences observed during this collection period, and thus represent real-world challenges encountered in textile production processes. The dataset also showcases normal images. We announce two folders, train and test in the same folder architecture,

0 papers0 benchmarksImages

AntM2C (Ant-Group Multi-Scenario Multi-Modal CTR dataset)

We release a large-scale Multi-Scenario Multi-Modal CTR dataset named AntM2C, built from real industrial data from Alipay. This dataset offers an impressive breadth and depth of information, covering CTR data from four diverse business scenarios, including advertisements, consumer coupons, mini-programs, and videos. Unlike existing datasets, AntM2C provides not only ID-based features but also five textual features and one image feature for both users and items, supporting more delicate multi-modal CTR prediction.

0 papers0 benchmarksImages, Tabular, Texts

ViMedical_Disease

This dataset contains over 12K+ questions and symptoms related to various common diseases in Vietnamese. It's designed to aid in the classification of medical symptoms and provide preliminary disease identification. The dataset covers a wide range of diseases, including cardiovascular, digestive, neurological, dermatological, endocrine, and others.

0 papers0 benchmarks

ViMirr

we construct the ViMirr dataset, which has 19,255 frames from 276 videos

0 papers0 benchmarks

InScope

Perception systems of autonomous vehicles are susceptible to occlusion, especially when examined from a vehicle-centric perspective. Such occlusion can lead to overlooked object detections, e.g., larger vehicles such as trucks or buses may create blind spots where cyclists or pedestrians could be obscured, accentuating the safety concerns associated with such perception system limitations. To mitigate these challenges, the vehicle-to-everything (V2X) paradigm suggests employing an infrastructure-side perception system (IPS) to complement autonomous vehicles with a broader perceptual scope. Nevertheless, the scarcity of real-world 3D infrastructure-side datasets constrains the advancement of V2X technologies. To bridge these gaps, this paper introduces a new 3D infrastructure-side collaborative perception dataset, abbreviated as inscope. Notably, InScope is the first dataset dedicated to addressing occlusion challenges by strategically deploying multiple-position Light Detection and Ran

0 papers0 benchmarks

MLHME-38K

Source: MLHME-38K

0 papers0 benchmarks

3DPotatoTwinDemo

Demo dataset with 5 partial and complete 3D shapes of potato tubers.

0 papers0 benchmarks

sheeep ruminate behavior

sheeep ruminate behavior image and label

0 papers0 benchmarks

Sheeep ruminate behavior

sheeep ruminate behavior image and label

0 papers0 benchmarks

DDoS-detect (Towards Resource-Efficient DDoS Detection in IoT: Leveraging Feature Engineering of System and Network Usage Metrics)

The Internet of Things (IoT) is omnipresent, exposing a large number of devices that often lack security controls to the public Internet. In the modern world, many everyday processes depend on these devices, and their service outage could lead to catastrophic consequences. There are many Deep Packet Inspection (DPI) based intrusion detection systems (IDS). However, their linear computational complexity induced by the event-driven nature poses a power-demanding obstacle in resource-constrained IoT environments. In this paper, we shift away from the traditional IDS as we introduce a novel and lightweight framework, relying on a time-driven algorithm to detect Distributed Denial of Service (DDoS) attacks by employing Machine Learning (ML) algorithms leveraging the newly engineered features containing system and network utilization information. These features are periodically generated, and there are only ten of them, resulting in a low and constant algorithmic complexity. Moreover, we lev

0 papers0 benchmarks

DDoS-detect: Towards Resource-Efficient DDoS Detection in IoT (Towards Resource-Efficient DDoS Detection in IoT: Leveraging Feature Engineering of System and Network Usage Metrics)

The Internet of Things (IoT) is omnipresent, exposing a large number of devices that often lack security controls to the public Internet. In the modern world, many everyday processes depend on these devices, and their service outage could lead to catastrophic consequences. There are many Deep Packet Inspection (DPI) based intrusion detection systems (IDS). However, their linear computational complexity induced by the event-driven nature poses a power-demanding obstacle in resource-constrained IoT environments. In this paper, we shift away from the traditional IDS as we introduce a novel and lightweight framework, relying on a time-driven algorithm to detect Distributed Denial of Service (DDoS) attacks by employing Machine Learning (ML) algorithms leveraging the newly engineered features containing system and network utilization information. These features are periodically generated, and there are only ten of them, resulting in a low and constant algorithmic complexity. Moreover, we lev

0 papers0 benchmarks

Towards Resource-Efficient DDoS Detection in IoT: Leveraging Feature Engineering of System and Network Usage Metrics

The Internet of Things (IoT) is omnipresent, exposing a large number of devices that often lack security controls to the public Internet. In the modern world, many everyday processes depend on these devices, and their service outage could lead to catastrophic consequences. There are many Deep Packet Inspection (DPI) based intrusion detection systems (IDS). However, their linear computational complexity induced by the event-driven nature poses a power-demanding obstacle in resource-constrained IoT environments. In this paper, we shift away from the traditional IDS as we introduce a novel and lightweight framework, relying on a time-driven algorithm to detect Distributed Denial of Service (DDoS) attacks by employing Machine Learning (ML) algorithms leveraging the newly engineered features containing system and network utilization information. These features are periodically generated, and there are only ten of them, resulting in a low and constant algorithmic complexity. Moreover, we lev

0 papers0 benchmarks

CHAINS

The CHAINS dataset is utilized to evaluate the proposed system. This dataset includes both normal speech, referred to as solo, and whispered speech, known as whsp. The dataset comprises recordings from 36 speakers, dominantly with Irish dialects. To develop a model capable of accurately capturing the acoustic features of the Irish dialect, the normal speech portion of this dataset is used for system training, while the whispered speech portion is employed for evaluation.

0 papers0 benchmarks

Digital Forensics 2023 dataset - DF2023

The deliberate manipulation of public opinion, especially through altered images, poses a significant danger to society. To fight this issue on a technical level we support the research community by releasing the Digital Forensics 2023 (DF2023) training and validation dataset.

0 papers0 benchmarksImages
PreviousPage 663 of 1000Next