19,997 machine learning datasets
19,997 dataset results
Existing benchmark datasets in real-world distribution shifts are generally synthetically generated via augmentations to simulate real-world shifts such as weather and camera rotation. The UCF101-DS dataset consists of real-world distribution shifts from user-generated videos without synthetic augmentation. It has videos for 47 UCF-101 classes with 63 different distribution shifts that can be categorized into 15 categories. A total of 536 unique videos split into a total of 4,708 clips. Each clip ranges from 7 to 10 seconds long.
LibriS2S is a Speech to Speech Translation (S2ST) dataset build further upon existing resources. The dataset provides English-German speech and text quadruplets ranging just over 50 hours for both languages.
This dataset contains 304 manual evaluations of class-level software maintainability, drawn from 5 open-source projects: ArgoUML, Art of Illusion, Diary Management, JUnit 4, JSweet. Each Java class is labelled along 5 axis: readability, understandability, complexity, modularity and overall maintainability. Each Java class was assessed by several experts independently of its relation to other classes.
4 different synthetic datasets generated by Blender
The MIMIC-IV-ICD9 dataset, featuring the top 50 most frequently occurring labels.
The MIMIC-IV-ICD10 dataset, featuring the top 50 most frequently occurring labels.
Spectrum data on CH4/Air flame emission with 200ms exposure and 2s exposure
The buildingSMART Data Dictionary (bSDD) is an online service that hosts classifications and their properties, allowed values, units and translations. The bSDD allows linking between all the content inside the database. It provides a standardized workflow to guarantee data quality and information consistency.
ChatLog is a coarse-to-fine temporal dataset called ChatLog, consisting of two parts that update monthly and daily:
arxiv : https://arxiv.org/abs/2304.11708
a dataset of reading pointer meter
The SF100 corpus of classes is a statistically representative sample of 100 Java projects from SourceForge, which is a popular open source repository (more than 300,000 projects with more than two million registered users). Because SourceForge is home to many old and stale projects, we have extended SF100 with the 10 most popular projects, resulting in a revised corpus of classes, SF110.
The dataset contains 256x256 tiles extracted from Whole Slide Images (WSI) of mouse liver tissue stained with H&E and Masson's Trichrome. WSIs were acquired with a Zeiss AxioScan scanner with a 20× objective at a resolution of 0.221 µm/pix and subsequently subsampled with a factor of 1:2, which resulted in a 0.442 µm/pixel resolution.
ChCatExt is composed of BidAnn (bid announcement), FinAnn (financial announcement) and CreRat (credit rating report). It is designed for re-construct catalog trees from documents.
The dataset is composed of 95 unique document texts spanning the period 2005-2022. This dataset makes available a corpus of documentary sources useful for outlining case studies related to scenarios in which the DPO finds himself operating in the performance of his daily activities.
DP0E is a public dataset of anti-counterfeiting printable graphical codes (PGC) based on DataMatrix modulation.
SLATS is a dataset which covers two data domains. Each domain is populated by a variant of a LArTPC detector simulation used in the ProtoDUNE-SP experiment. The two domains differ in one feature—the detector response function. The domain real is generated with a 2D response, and the fake domain is generated with a quasi-1D response. The dataset can be used to train unpaired image translation algorithms.
Multimedia Goal-oriented Generative Script Learning Dataset This link contains a dataset consisting of multimedia steps for two categories: gardening and crafts. The dataset consists of a total of 79,089 multimedia steps across 5,652 tasks.
The Human Phenotype Ontology (HPO) graph is a standardized vocabulary of human phenotypic abnormalities and their relationships. It represents these abnormalities as nodes in a graph, with edges indicating relationships such as subtypes or overlapping features. The HPO graph is organized in a hierarchical structure, with more general terms at the top and more specific terms at the bottom. The ontology provides a framework for the annotation of human genetic variations, aiding in the diagnosis of rare genetic disorders and the identification of potential therapeutic targets.
A large film style dataset