19,997 machine learning datasets
19,997 dataset results
This is our replication package for our study on Benchmarking scalability of stream processing frameworks deployed as microservices in the cloud.
The Multi-Layer Materials Science corpus (MuLMS) consists of 50 documents (licensed CC BY) from the materials science domain, spanning across the following 7 subareas: "Electrolysis", "Graphene", "Polymer Electrolyte Fuel Cell (PEMFC)", "Solid Oxide Fuel Cell (SOFC)", "Polymers", "Semiconductors" and "Steel". It was exhaustively annotated by domain experts. There are annotations on sentence-level and token-level for the following NLP tasks: measurement frame detection, NER, relation extraction, and argumentative zones classifications.
This paper analyses two hitherto unstudied sites sharing state-backed disinformation, Reliable Recent News (rrn.world) and WarOnFakes (waronfakes.com), which publish content in Arabic, Chinese, English, French, German, and Spanish.
CovidET-Appraisals is the most comprehensive dataset to-date that assesses 24 cognitive appraisal dimensions of emotions, each with a natural language rationale, across 241 Reddit posts. CovidET-Appraisals presents an ideal testbed to evaluate the ability of large language models — excelling at a wide range of NLP tasks — to automatically assess and explain cognitive appraisals.
The dataset contains 50 hours for high quality speech samples from a native speaker and 2 more hours of lower quality recordings from a different speaker
Generated using the script below: https://github.com/zenineasa/MasterThesis/blob/main/Code/dataGenerator.py
This dataset is a multi-labelled SMILES odor dataset with 138 odor descriptors. This dataset was created for replicating the paper: A principal odor map unifies diverse tasks in olfactory perception.
Collection of news websites in low-resource languages.
StoryBooks for 174 unique languages.
PETA: Evaluating the Impact of Protein Transfer Learning with Sub-word Tokenization on Downstream Applications
WEATHub is a dataset containing 24 languages. It contains words organized into groups of (target1, target2, attribute1, attribute2) to measure the association target1:target2 :: attribute1:attribute2. For example target1 can be insects, target2 can be flowers. And we might be trying to measure whether we find insects or flowers pleasant or unpleasant. The measurement of word associations is quantified using the WEAT metric in our paper. It is a metric that calculates an effect size (Cohen's d) and also provides a p-value (to measure statistical significance of the results). In our paper, we use word embeddings from language models to perform these tests and understand biased associations in language models across different languages.
This dataset, adapted from COCO Caption, is designed for the Image Caption task and evaluates multimodal model editing in terms of reliability, stability and generality. You can download the dataset from here
This dataset, adapted from VQAv2, is designed for the Visual Question Answering task and evaluates multimodal model editing in terms of reliability, stability and generality. You can download the dataset from here
The corpus contains review sentences mostly of products in electronics domain, annotated and segregated into 4 comparison categories. Each comparison sentence is annotated with names of the products (PROD1 and PROD2), the aspect (ASP) and the predicate (PRED). Dataset contains sentences after auto-labeling on SNAP dataset and manually labeled sentences from the following corpora:
RVL-CDIP_MP is our first contribution to retrieve the original documents of the IIT-CDIP test collection which were used to create RVL-CDIP. Some PDFs or encoded images were corrupt, which explains that we have around 500 fewer instances. By leveraging metadata from OCR-IDL , we matched the original identifiers from IIT-CDIP and retrieved them from IDL using a conversion.
RVL-CDIP_MP-N can serve its original goal as a covariate shift test set, now for multi-page document classification. We were able to retrieve the original full documents from DocumentCloud and Web Search.
CiNAT Birds 2021 (Cross-View iNaturalist-2021 Birds) dataset contains ground-level images of bird species along with satellite images associated with the geolocation of the ground-level images. In total, there are 413,959 pairs for training and 14,831 pairs for validation and testing. The ground-level images are of varying sizes while the satellite images are of size 256x256. Additionally, the dataset comes with rich metadata for each image - geolocation, date, observer id, taxonomy.
We introduce a large semi-automatically generated dataset of ~400,000 descriptive sentences about commonsense knowledge that can be true or false in which negation is present in about 2/3 of the corpus in different forms that we use to evaluate LLMs.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
This is the Infrared Elephant Images Dataset (named 'EleThermal dataset') collected from here and annotated by our project, released under GPLv3. Therefore, if you use the annotated 'EleThermal' dataset for any research or other product by any means, please acknowledge the following two works by citing them.