19,997 machine learning datasets
19,997 dataset results
The DeepSpeak dataset contains over 43 hours of real and deepfake footage of people talking and gesturing in front of their webcams. The source data was collected from a diverse set of participants in their natural environments and the deepfakes were generated using state-of-the-art open-source lip-sync and face-swap software.
A outdoor dataset for UGNA-VPR
Data files with the information required to replicate all the experiments reported in the paper:
Introduced by Khan. et al. Divide and conquer: Ill-light image enhancement via hybrid deep network https://www.sciencedirect.com/science/article/abs/pii/S0957417421004759
This is the list of datasets used for Conti's FNS-Funded projects
Key Points
InpaintCOCO is a benchmark to understand fine-grained concepts in multimodal models (vision-language) similar to Winoground. To our knowledge InpaintCOCO is the first benchmark, which consists of image pairs with minimum differences, so that the visual representation can be analyzed in a more standardized setting.
POPCORN is a French dataset consisting of 400 validation texts and 400 training texts, all written and annotated manually. The texts are concise and factual, resembling information reports. The annotations, based on the ontology described below, allow for the training and evaluation of models in Information Extraction tasks, including Named Entity Recognition, Coreference Resolution, and Relation Extraction.
Scene Recognition is a problem, where a set of visible objects must be correctly associated with objects marked on a semantic map - this problem is also sometimes called a Data Association. Please note, that Scene Recognition in terms where the observed scene must be labeled in terms such as 'kitchen', 'bedroom' and so on is a different problem.
List of ontologies in the domain of Materials Science and Engineering.
Heel Bone X-Ray Dataset consists of 3,956 X-ray images of the foot, primarily focused on detecting and classifying heel bone diseases. The images were obtained from Kirkuk General Hospital in Digital Imaging and Communications in Medicine (DICOM) format and converted to JPG format using the MicroDicom tool.
Machine Learning for Two-Sample Testing under Right-Censored Data: A Simulation Study
Dataset: RGB-D Images for Real-World and Synthetic Object Scenes This dataset consists of both real-world and synthetic RGB-D images, designed for object detection, classification, and segmentation tasks, particularly for primitive shape recognition.
Medical report generation (MRG), which aims to automatically generate a textual description of a specific medical image (e.g., a chest X-ray), has recently received increasing research interest. Building on the success of image captioning, MRG has become achievable. However, generating language-specific radiology reports poses a challenge for data-driven models due to their reliance on paired image-report chest X-ray datasets, which are labor-intensive, time-consuming, and costly. In this paper, we introduce a chest X-ray benchmark dataset, namely CASIA-CXR, consisting of high-resolution chest radiographs accompanied by narrative reports originally written in French. To the best of our knowledge, this is the first public chest radiograph dataset with medical reports in this particular language. Importantly, we propose a simple yet effective multimodal encoder-decoder contextually-guided framework for medical report generation in French. We validated our framework through intra-language
Overview The IITKGP_Fence dataset is designed for tasks related to fence-like occlusion detection, defocus blur, depth mapping, and object segmentation. The captured data vaies in scene composition, background defocus, and object occlusions. The dataset comprises both labeled and unlabeled data, as well as additional video and RGB-D data. The contains ground truth occlusion masks (GT) for the corresponding images. We created the ground truth occlusion labels in a semi-automatic way with user interaction.
1222
We introduce an annotated dataset of five thousand human labeled pareidolic face images, called ``Faces in Things''. Faces in Things is derived from the LAION-5B dataset and annotated for key face attributes and bounding boxes
Post-Spraying Image Evaluation This dataset is for the paper Deep Learning for Precision Agriculture: Post-Spraying Evaluation and Deposition Estimation (https://arxiv.org/abs/2409.16213).
CodeSCAN is the first large-scale and diverse dataset of coding screenshots with pixel-perfect annotations. It features:
test-dataset