TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

MIM-GOLD-NER

The MIM-GOLD-NER dataset is an Icelandic named entity (NE) corpus. It is a version of the MIM-GOLD corpus that has been specifically tagged for named entities. In this dataset, over 48,000 NEs (named entities) are labeled within a corpus of one million tokens. Researchers and developers can use this dataset to train named entity recognizers for Icelandic¹²³.

0 papers0 benchmarks

SweDN

The SweDN 1.0 dataset is a valuable resource for natural language processing (NLP) tasks, specifically text summarization. Let's delve into the details:

0 papers0 benchmarks

DMS (Dense Material Segmentation Dataset)

The Dense Material Segmentation Dataset (DMS) consists of 3 million polygon labels of material categories (metal, wood, glass, etc) for 44 thousand RGB images. The dataset is described in the research paper, A Dense Material Segmentation Dataset for Indoor and Outdoor Scene Parsing.

0 papers0 benchmarksImages

MVTec ITODD

The MVTec Industrial 3D Object Detection Dataset (MVTec ITODD), introduced by Bertram Drost, Markus Ulrich, Paul Bergmann, and Carsten Steger from MVTec Software GmbH, is a valuable resource for 3D object detection and pose estimation in industrial contexts¹²³. Here are the key details about this dataset:

0 papers0 benchmarks

IC-BIN

The IC-BIN dataset was introduced by Doumanoglou et al. as part of their research on recovering 6D object pose and predicting next-best-view in the crowd¹². This dataset is specifically designed to address the challenges posed by reflective objects in robotic bin-picking scenarios.

0 papers0 benchmarks

IC-MI

The IC-MI dataset, introduced by Tejani et al., is part of the Benchmark for 6D Object Pose Estimation (BOP). Let's delve into the details:

0 papers0 benchmarks

MIMIC Meme Dataset (Misogyny Identification in Multimodal Internet Content in Hindi-English Code-Mix Language)

This dataset endeavors to fill the research void by presenting a meticulously curated collection of misogynistic memes in a code-mixed language of Hindi and English. It introduces two sub-tasks: the first entails a binary classification to determine the presence of misogyny in a meme, while the second task involves categorizing the misogynistic memes into multiple labels, including Objectification, Prejudice, and Humiliation.

0 papers0 benchmarksImages, Texts

tomato detection (A dataset of tomato fruits images for object detection in the complex lighting environment of plant factories)

Plant factories are an advanced form of facility agriculture that enable efficient plant cultivation through controllable environmental conditions, making them highly suitable for the automation and intelligent application of machinery. Tomato cultivation in plant factories has significant economic and agricultural value and can be utilized for various applications such as seedling cultivation, breeding, and genetic engineering. However, manual completion is still required for operations such as detection, counting, and classification of tomato fruits, and the application of machine detection is currently inefficient. Furthermore, research on the automation of tomato harvesting in plant factory environments is limited due to the lack of a suitable dataset. To address this issue, a tomato fruit dataset was constructed for plant factory environments, named as TomatoPlantfactoryDataset, which can be quickly applied to multiple tasks, including the detection of control systems, harvesting

0 papers0 benchmarksImages

tomato fruits detection (A dataset of tomato fruits images for object detection in the complex lighting environment of plant factories)

Plant factories are an advanced form of facility agriculture that enable efficient plant cultivation through controllable environmental conditions, making them highly suitable for the automation and intelligent application of machinery. Tomato cultivation in plant factories has significant economic and agricultural value and can be utilized for various applications such as seedling cultivation, breeding, and genetic engineering. However, manual completion is still required for operations such as detection, counting, and classification of tomato fruits, and the application of machine detection is currently inefficient. Furthermore, research on the automation of tomato harvesting in plant factory environments is limited due to the lack of a suitable dataset. To address this issue, a tomato fruit dataset was constructed for plant factory environments, named as TomatoPlantfactoryDataset, which can be quickly applied to multiple tasks, including the detection of control systems, harvesting

0 papers0 benchmarksImages

Data on: Cell Signaling and Targeted Drug Therapy of Cancer

The landmark Cancer Genomics Program launched in 2006 has contributed immensely to the awareness of the importance of cancer genomics in our understanding of cancer over the past decade and has begun to change the way the disease is treated in clinic. A large number of mutations contribute to cancer and predicting the effects of mutations using in silico tools has become a frequently used approach, but the use of next-generation sequencing-based approaches in clinical diagnosis has also led to a considerable increase in data and a vast number of variants of uncertain significance that require further analysis and validation to achieve the development goals. These data cannot be analyzed simply by using the tools and techniques traditionally available to better understand the origin and evolution of cancer and therefore to achieve this goal, a cancer reference framework through modeling of genome sequencing data has been proposed for the systematic identification of representative drive

0 papers0 benchmarks

ChineseSquad

ChineseSquad (中文机器阅读理解数据集) is a dataset specifically designed for Chinese machine reading comprehension. It is created by translating and manually correcting the original SQuAD (Stanford Question Answering Dataset) into Chinese. The dataset includes both V1.1 and V2.0 versions of SQuAD. However, due to some translation challenges (especially with short answers and document translations), the Chinese version has slightly fewer examples compared to the original English SQuAD¹.

0 papers0 benchmarks

Redteaming Resistance Benchmark

The Redteaming Resistance Benchmark is a project aimed at evaluating the robustness of language models, both open-source and black-box, through redteaming attacks. These attacks involve systematically challenging and testing models with carefully crafted prompts to uncover their failure modes and vulnerabilities. In other words, it reveals where these models are susceptible to generating problematic outputs¹².

0 papers0 benchmarks

Student-Teacher Prompting

Student-Teacher Prompting is an instructional strategy used to guide a learner's behavior. It is particularly helpful for teaching new skills or encouraging desired behaviors. Here are some examples of different types of prompts that teachers or educators might use:

0 papers0 benchmarks

AlexMI MOABB (Alex Motor Imagery dataset.)

0 papers0 benchmarks

BI2012 MOABB (P300 dataset BI2012 from a "Brain Invaders" experiment.)

0 papers0 benchmarks

BI2013a MOABB (P300 dataset BI2013a from a "Brain Invaders" experiment.)

0 papers0 benchmarks

BI2014a MOABB (P300 dataset BI2014a from a "Brain Invaders" experiment.)

0 papers0 benchmarks

BI2014b MOABB (P300 dataset BI2014b from a "Brain Invaders" experiment.)

0 papers0 benchmarks

BI2015a MOABB (P300 dataset BI2015a from a "Brain Invaders" experiment.)

0 papers0 benchmarks

BI2015b MOABB (P300 dataset BI2015b from a "Brain Invaders" experiment.)

0 papers0 benchmarks
PreviousPage 659 of 1000Next