19,997 machine learning datasets
19,997 dataset results
This dataset consists of the Graphcast model checkpoints produced during the fine-tuning process of (Subich 2024).
Please refer this paper
The Deepfake face detection task involves a facial image of unknown authenticity for testing. While most deepfake detection methods take only the image as input, our literature demonstrates that conditioning the deepfake detector on identity—i.e., knowing whose deepfake face the picture might be—can enhance detection performance. Existing deepfake detection datasets, such as FaceForensics++ and DFDC, do not include identity information for authentic and deepfake faces. This dataset contains facial images of 45 specific individuals, divided into train and test sets, including a total of 23k authentic and 22k deepfake images. Having a specific individual's images in both the train and test sets allows us to assess detection performance for that individual. The dataset is curated so that the train and test sets are from two independent sources. The train images are curated from the CelebDFv2 dataset, and the test images are curated from the CACD dataset. Deepfake faces are generated using
A large dataset of around 40000 Reddit posts was collected from r/suicidewatch and other non-suicidal subreddits. The posts collected from r/suicidewatch are annotated with suicidal and other posts collected from a variety of groups like r/sports, r/anxiety, r/politics, and more are annotated with non-suicidal. Then this dataset has been used to feed various advanced deep-learning models to report a comparative evaluation of these models.
A pair Deblurring Benchmarking Dataset
Karrierewege Dataset
Karrierewege+ Dataset
The largest video inpainting dataset comprises over 390K clips (> 866.7 hours), featuring precise masks and detailed video captions.
The benchmark for VPData, the largest video inpainting dataset, which comprises over 390K clips (> 866.7 hours) and features precise masks and detailed video captions.
It is a large-scale multimodal patent dataset with detailed captions for design patent figures.
Simulations
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
This dataset consists of annotated images and videos of smoke resulting from prescribed burning events in Finnish boreal forests. The dataset was created to train and validate learning-based methods for wildfire detection and smoke segmentation and its effectiveness in doing so was shown in the linked studies.
This archived Paleoclimatology Study is available from the NOAA National Centers for Environmental Information (NCEI), under the World Data Service (WDS) for Paleoclimatology. The associated NCEI study type is Paleoceanography. The data include parameters of paleoceanography with a geographic location of Eastern Pacific Ocean. The time period coverage is from 15190 to 1330 in calendar years before present (BP).
This dataset contains Material-Point-Method (MPM) simulations for various materials, including water, sand, plasticine, elasticity, jelly, rigid collisions, and melting. Each material is represented as point-clouds that evolve over time. The dataset is designed for learning and predicting MPM-based physical simulations.
The following datasets:
Replication Material This document contains the necessary materials and instructions to replicate the findings presented in our paper. We provide comprehensive information on the data sources, code, and analytical procedures used in our study. The replication package includes raw data files, data cleaning scripts, and analysis code. We encourage users to contact us with any questions or issues encountered during the replication process.
A 3D skeleton-based group activity understanding dataset.