19,997 machine learning datasets
19,997 dataset results
The MIM-GOLD-NER dataset is an Icelandic named entity (NE) corpus. It is a version of the MIM-GOLD corpus that has been specifically tagged for named entities. In this dataset, over 48,000 NEs (named entities) are labeled within a corpus of one million tokens. Researchers and developers can use this dataset to train named entity recognizers for Icelandic¹²³.
The SweDN 1.0 dataset is a valuable resource for natural language processing (NLP) tasks, specifically text summarization. Let's delve into the details:
The Dense Material Segmentation Dataset (DMS) consists of 3 million polygon labels of material categories (metal, wood, glass, etc) for 44 thousand RGB images. The dataset is described in the research paper, A Dense Material Segmentation Dataset for Indoor and Outdoor Scene Parsing.
The MVTec Industrial 3D Object Detection Dataset (MVTec ITODD), introduced by Bertram Drost, Markus Ulrich, Paul Bergmann, and Carsten Steger from MVTec Software GmbH, is a valuable resource for 3D object detection and pose estimation in industrial contexts¹²³. Here are the key details about this dataset:
The IC-BIN dataset was introduced by Doumanoglou et al. as part of their research on recovering 6D object pose and predicting next-best-view in the crowd¹². This dataset is specifically designed to address the challenges posed by reflective objects in robotic bin-picking scenarios.
The IC-MI dataset, introduced by Tejani et al., is part of the Benchmark for 6D Object Pose Estimation (BOP). Let's delve into the details:
This dataset endeavors to fill the research void by presenting a meticulously curated collection of misogynistic memes in a code-mixed language of Hindi and English. It introduces two sub-tasks: the first entails a binary classification to determine the presence of misogyny in a meme, while the second task involves categorizing the misogynistic memes into multiple labels, including Objectification, Prejudice, and Humiliation.
Plant factories are an advanced form of facility agriculture that enable efficient plant cultivation through controllable environmental conditions, making them highly suitable for the automation and intelligent application of machinery. Tomato cultivation in plant factories has significant economic and agricultural value and can be utilized for various applications such as seedling cultivation, breeding, and genetic engineering. However, manual completion is still required for operations such as detection, counting, and classification of tomato fruits, and the application of machine detection is currently inefficient. Furthermore, research on the automation of tomato harvesting in plant factory environments is limited due to the lack of a suitable dataset. To address this issue, a tomato fruit dataset was constructed for plant factory environments, named as TomatoPlantfactoryDataset, which can be quickly applied to multiple tasks, including the detection of control systems, harvesting
Plant factories are an advanced form of facility agriculture that enable efficient plant cultivation through controllable environmental conditions, making them highly suitable for the automation and intelligent application of machinery. Tomato cultivation in plant factories has significant economic and agricultural value and can be utilized for various applications such as seedling cultivation, breeding, and genetic engineering. However, manual completion is still required for operations such as detection, counting, and classification of tomato fruits, and the application of machine detection is currently inefficient. Furthermore, research on the automation of tomato harvesting in plant factory environments is limited due to the lack of a suitable dataset. To address this issue, a tomato fruit dataset was constructed for plant factory environments, named as TomatoPlantfactoryDataset, which can be quickly applied to multiple tasks, including the detection of control systems, harvesting
The landmark Cancer Genomics Program launched in 2006 has contributed immensely to the awareness of the importance of cancer genomics in our understanding of cancer over the past decade and has begun to change the way the disease is treated in clinic. A large number of mutations contribute to cancer and predicting the effects of mutations using in silico tools has become a frequently used approach, but the use of next-generation sequencing-based approaches in clinical diagnosis has also led to a considerable increase in data and a vast number of variants of uncertain significance that require further analysis and validation to achieve the development goals. These data cannot be analyzed simply by using the tools and techniques traditionally available to better understand the origin and evolution of cancer and therefore to achieve this goal, a cancer reference framework through modeling of genome sequencing data has been proposed for the systematic identification of representative drive
ChineseSquad (中文机器阅读理解数据集) is a dataset specifically designed for Chinese machine reading comprehension. It is created by translating and manually correcting the original SQuAD (Stanford Question Answering Dataset) into Chinese. The dataset includes both V1.1 and V2.0 versions of SQuAD. However, due to some translation challenges (especially with short answers and document translations), the Chinese version has slightly fewer examples compared to the original English SQuAD¹.
The Redteaming Resistance Benchmark is a project aimed at evaluating the robustness of language models, both open-source and black-box, through redteaming attacks. These attacks involve systematically challenging and testing models with carefully crafted prompts to uncover their failure modes and vulnerabilities. In other words, it reveals where these models are susceptible to generating problematic outputs¹².
Student-Teacher Prompting is an instructional strategy used to guide a learner's behavior. It is particularly helpful for teaching new skills or encouraging desired behaviors. Here are some examples of different types of prompts that teachers or educators might use: