TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

PlantSeg

We established a large-scale plant disease segmentation dataset named PlantSeg. PlantSeg comprises more than 11,400 images of 115 different plant diseases from various environments, each annotated with its corresponding segmentation label for diseased parts. To the best of our knowledge, PlantSeg is the largest plant disease segmentation dataset containing in-the-wild images. Our dataset enables researchers to evaluate their models and provides a valid foundation for the development and benchmarking of plant disease segmentation algorithms.

1 papers0 benchmarksImages

AdvSuffixes (Adversarial Suffixes)

AdvSuffixes - Information AdvSuffixes is a curated dataset of adversarial prompts and suffixes designed to evaluate and enhance the robustness of large language models (LLMs) against adversarial attacks. By appending these suffixes to standard prompts, researchers and developers can explore and analyze how LLMs respond to potentially harmful input scenarios. This dataset is heavily inspired by AdvBench.

1 papers0 benchmarksTexts

approved_drug_target (Approved Drug SMILES and Protein Sequence Dataset)

This dataset provides a curated collection of approved drug Simplified Molecular Input Line Entry System (SMILES) strings and their associated protein sequences. Each small molecule has been approved by at least one regulatory body, ensuring the safety and relevance of the data for computational applications. The dataset includes 1,660 approved small molecules and their 2,093 related protein targets.

1 papers0 benchmarksTexts

Russian Sentences POS tagged

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksTexts

DeVAn (Dense Video Annotation for Video-Language Models)

DeVAn is a multi-modal dataset containing 8.5K video clips carefully selected from previously published YouTube-based video datasets (YouTube-8M and YT-Temporal-1B) that integrate visual and auditory information. Over the span of 10 months, a team of 24 human annotators (college and graduate level students) created 5 short captions (1 sentence each) and 5 long summaries (3-10 sentences) for each video clip, resulting in a rich and comprehensive human-annotated dataset that serves as a robust ground truth for subsequent model training and evaluation.

1 papers0 benchmarksVideos

Lunar Navigation

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

Lunar Simulation

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

Park with Pavilion

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

Speed Limit Signs

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

https://www.aminer.cn/cosnet

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

https://networkrepository.com

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

Khayyam Offline Persian Handwriting Dataset

Handwriting analysis is still an important application in machine learning. A basic requirement for any handwriting recognition application is the availability of comprehensive datasets. Standard labelled datasets play a significant role in training and evaluating learning algorithms. In this paper, we present the Khayyam dataset as another large unconstrained handwriting dataset for elements (words, sentences, letters, digits) of the Persian language. We intentionally concentrated on collecting Persian word samples which are rare in the currently available datasets. Khayyam's dataset contains 44000 words, 60000 letters, and 6000 digits. Moreover, the forms were filled out by 400 native Persian writers. To show the applicability of the dataset, machine learning algorithms are trained on the digits, letters, and word data and results are reported. This dataset is available for research and academic use.

1 papers0 benchmarks

Mon(IoT)r Testbed

Mon(IoT)r Testbed The Mon(IoT)r Testbed is the traffic capture software developed for the Mon(IoT)r Lab. It is currently deployed at Imperial College London and is being installed at Politecnico di Torino.

1 papers0 benchmarks

Psi-Class Intersection Numbers

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

WiFiCam

WiFiCam dataset for through-wall imaging based on WiFi channel state information. The corresponding source code repository is located at: https://github.com/StrohmayerJ/wificam

1 papers0 benchmarksEnvironment, Images, RGB Video, Time series

SumIPCC

The dataset contains 140 paragraphs from climate change reports with associated aspect-based (i.e. query-focused) summaries, that were produced by experts especially for policy-makers.

1 papers0 benchmarksTexts

TaxaBench-8k

A truly multimodal dataset for benchmarking deep learning models on ecological tasks.

1 papers0 benchmarks

iSatNat

A dataset consisting of paired ground-level images of species in the inat-2021 dataset and corresponding satellite imagery. The dataset has approximately 2.7M pairs.

1 papers0 benchmarks

3D Flow Shapes

The dataset consists of high-resolution three-dimensional (3D) turbulent flow simulations. It captures intricate vortex structures caused by a variety of shapes within a channel flow environment. The dataset is generated using OpenFOAM in large eddy simulation (LES) mode, ensuring the preservation of detailed turbulent characteristics across all spatial scales.

1 papers0 benchmarksTime series, Videos

noise2image dataset (Noise2Image dataset for event cameras)

Noise event recordings with static scenes. The static scenes are played on a display that does not flicker, so all events are not triggered by brightness changes and are pure noise. The goal for our noise2image paper (https://arxiv.org/abs/2404.01298) is to recover the static scene from the noise event recording using the noise-to-intensity dependency.

1 papers0 benchmarks
PreviousPage 530 of 1000Next