19,997 machine learning datasets
19,997 dataset results
A dataset for Image-Goal Navigation in Habitat based on Gibson scenes.
The acquisition over the VIS and TIR data was performed by a commercial thermal camera Testo 882-3. We have used a second external camera to obtain the NIR data. In this case we have built a NIR camera using a webcam changing the default optical filter for a couple of Kodak filters for IR. We have also used a printed circuit board with 16 infra-red LEDs that provide the infra-red illumination.
Original images and images with RUSTICO filters applied
The French national meteorological service published an open-access dataset of hourly weather observations in Brittany, France, for the month of January 2014. In addition to the graph of ground weather stations, the dataset contains hourly readings of those stations. Readings include temperatures, wind characteristics, rain, and other information.
Te NVALT-8 study (m=200 participants) examined if nadroparin combined with chemotherapy could reduce cancer relapse after surgical removal of a non-small cell lung tumour.
The NVALT-11 study considered the effect of profylactic brain radiation versus observation in ($m$=174) patients with advanced non-small cell lung cancer.
Is there any correlation between the impact of a scientific conference and the venue where it takes place? It seems that no one has tackled this issue before, so we decided to explore the possible implications. From the one hand, we considered the number of citations as indicator of the impact of a conference; from the other hand, we considered specific touristic indexes that characterize the venue. In this work we report on the results of the large scale analysis we conducted on the bibliographic data we extracted from nearly 4000 conference series in the Computer Science area and over 2.5 million papers spanning more than 30 years of research. Interestingly, we found out that the two aspects are indeed related and this is shown by the detailed analysis of the data.
Please see code repository. https://github.com/nlandolfi/acc2022treelinearcascades_stocks
Twitter dataset related to flood events onsets in Thailand and Nepal, focused on September 26/27, 2022, June 16/17 2021 and July 01/02 2021. The dataset has been processed with a VisualCit pipeline in order to automatically filter a relevant subset of posts through automated image analysis, using deep learning techniques. The posts were then geolocated using the CIME algorithm. Additional information about the data collection and data processing are described in http://arxiv.org/abs/2202.12014
Dataset Description This dataset contains +10k math memes. Memes were approved by admins before being shared with group members. Thus all memes follow the community standards. Memes are about college math or above.
EmoSpeech contains keywords with diverse emotions and background sounds, presented to explore new challenges in audio analysis.
A Zero-Shot Sketch-based Inter-Modal Object Retrieval Scheme for Remote Sensing Images
Sign Language Datasets for French Belgian Sign Language This dataset is built upon the work of Belgian linguists from the University of Namur. During eight years, they've collected and annotated 50 hours of videos depicting sign language conversation. 100 signers were recorded, making it one of the most representative sign language corpus. The annotation has been sanitized and enriched with metadata to construct two, easy to use, datasets for sign language recognition. One for continuous sign language recognition and the other for isolated sign recognition.
Dubbed series are gaining a lot of popularity in recent years with strong support from major media service providers. Such popularity is fueled by studies that showed that dubbed versions of TV shows are more popular than their subtitled equivalents.
The pretrained models from four image translation algorithms: ACL-GAN, Council-GAN, CycleGAN, and U-GAT-IT on three benchmarking datasets: Selfie2Anime, CelebA_gender, CelebA_glasses.
This is a synthetic dataset containing full images (instead of only cropped faces) that provides ground truth 3D gaze directions for multiple people in one image.
Synthetic training set: This set is constructed in the following two steps and will be used for estimation/training purposes. i) 84,000 275 pixel x 400 pixel ground-truth fingerprint images without any noise or scratches, but with random transformations (at most five pixels translation and +/-10 degrees rotation) were generated by using the software Anguli: Synthetic Fingerprint Generator. ii) 84,000 275 pixel x 400 pixel degraded fingerprint images were generated by applying random artifacts (blur, brightness, contrast, elastic transformation, occlusion, scratch, resolution, rotation) and backgrounds to the ground-truth fingerprint images. In total, it contains 168,000 fingerprint images (84,000 fingerprints, and two impressions - one ground-truth and one degraded - per fingerprint).
The standard evaluation protocol of Cross-View Time dataset allows for certain cameras to be shared between training and testing sets. This protocol can emulate scenarios in which we need to verify the authenticity of images from a particular set of devices and locations. Considering the ubiquity of surveillance systems (CCTV) nowadays, this is a common scenario, especially for big cities and high visibility events (e.g., protests, musical concerts, terrorist attempts, sports events). In such cases, we can leverage the availability of historical photographs of that device and collect additional images from previous days, months, and years. This would allow the model to better capture the particularities of how time influences the appearance of that specific place, probably leading to a better verification accuracy. However, there might be cases in which data is originated from heterogeneous sources, such as social media. In this sense, it is essential that models are optimized on camer
We collaborate with Blue Hexagon to release a dataset containing timestamped malware samples and well-curated family information for research purposes. The BODMAS dataset contains 57,293 malware samples and 77,142 benign samples collected from August 2019 to September 2020, with carefully curated family information (581 families). We also provide preprocessed feature vectors and metadata available to everyone. The malware binaries can be obtained per request.
Samples from NASA Perseverance and set of GAN generated synthetic images from Neural Mars.