FFT-75

CC BY 4.0Introduced 2019-08-16

The FFT-75 dataset contains randomly sampled, potentially overlapping file fragments from 75 popular file types. It is a diverse and balanced dataset which is labeled with class IDs and is ready for training supervised machine learning models. We distinguish 6 different scenarios with different granularity and provide variants with 512 and 4096-byte blocks. In each case, we sampled a balanced dataset and split the data as follows: 80% for training, 10% for testing and 10% for validation.