TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Colorectal-Liver-Metastases (Colorectal-Liver-Metastases | Preoperative CT and Survival Data for Patients Undergoing Resection of Colorectal Liver Metastases)

This collection consists of DICOM images and DICOM Segmentation Objects (DSOs) for 197 patients with Colorectal Liver Metastases (CRLM). The collection consists of a large, single-institution consecutive series of patients that underwent resection of CRLM and matched preoperative computed tomography (CT) scans for quantitative image analysis. Inclusion criteria were (a) pathologically confirmed resected CRLM, (b) available data from pathologic analysis of the underlying non-tumoral liver parenchyma and hepatic tumor, (c) available preoperative conventional portal venous contrast-enhanced multi-detector computed tomography (MDCT) performed within 6 weeks of hepatic resection. Patients with 90-day mortality or that had less than 24 months of follow-up were excluded. Additionally, because pathologic and radiographic alterations of the non-tumoral liver parenchyma caused by hepatic artery infusion (HAI) of chemotherapy are not well described, any patient who received preoperative HAI was e

0 papers0 benchmarks3D, Biomedical, Images, Medical

WTA/TLA (WTA/TLA: A UAV-captured Dataset for Semantic Segmentation of Energy Infrastructure)

WTA (Wind Turbine Aerial) and TLA (Transmission Line Aerial) are public datasets which contain a set of RGB images from wind turbine farms and transmission towers and power lines, along with semantic ground truth for relevant classes. This is the official repository of the paper: WTA/TLA: A UAV-captured Dataset for Semantic Segmentation of Energy Infrastructure (url).

0 papers0 benchmarksImages

UQ NIDS (ML-based NIDS NetFlow Datasets)

The datasets on this page are designed for machine learning-based Network Intrusion Detection Systems (NIDS) and are organised into the following high-level collections:

0 papers0 benchmarks

Beam-Level (5G) Time-Series Dataset

This dataset presents a novel, multi-variate time series specifically designed for advancing research in spatio-temporal forecasting. The primary goal of this dataset is to facilitate the accurate prediction of traffic throughput volumes across 5G communication networks.

0 papers0 benchmarksTime series

Urban Visual Pollution Dataset

<a href="https://gts.ai/dataset-download/urban-visual-pollution-dataset/" target="_blank">👉 Download the dataset here</a>

0 papers0 benchmarks

Mono kitti

Description:

0 papers0 benchmarks

CODrone

Applications of unmanned aerial vehicle (UAV) in logistics, agricultural automation, urban management, and emergency response are highly dependent on oriented object detection (OOD) to enhance visual perception. Although existing datasets for OOD in UAV provide valuable resources, they are often designed for specific downstream tasks. Consequently, they exhibit limited generalization performance in real flight scenarios and fail to thoroughly demonstrate algorithm effectiveness in practical environments. To bridge this critical gap, we introduce CODrone, a comprehensive oriented object detection dataset for UAVs that accurately reflects real-world conditions. It also serves as a new benchmark designed to align with downstream task requirements, ensuring greater applicability and robustness in UAV-based OOD. Based on application requirements, we identify four key limitations in current UAV OOD datasets-low image resolution, limited object categories, single-view imaging, and restricte

0 papers0 benchmarksImages

RFF DataSet

test

0 papers0 benchmarks

Video Dataset (Storytelling Video Dataset (Russian, Emotion, Gesture, Speech))

The Storytelling Video Dataset is a high-quality, human-reviewed multimodal dataset featuring over 700 full-body video recordings of native Russian speakers. Each video is 10+ minutes long and includes synchronized speech, facial expressions, gestures, and emotional variation. The dataset is ideal for research and development in:

0 papers0 benchmarksAudio, Speech, Texts, Videos

CelebDF-v2

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

0 papers0 benchmarks

ARF (Artificial Relationships in Fiction)

Artificial Relationships in Fiction Dataset Description Artificial Relationships in Fiction (ARF) is a synthetically annotated dataset for Relation Extraction (RE) in fiction, created from a curated selection of literary texts sourced from Project Gutenberg. The dataset captures the rich, implicit relationships within fictional narratives using a novel ontology and GPT-4o for annotation. ARF is the first large-scale RE resource designed specifically for literary texts, advancing both NLP model training and computational literary analysis.

0 papers0 benchmarksTexts

du lieu

gggg

0 papers0 benchmarks

Segmented Peripheral Blood Cells Using OpenCV

Description:

0 papers0 benchmarks

Jute Pest

Description:

0 papers0 benchmarks

PALMS

Data in this study come from western Ecuador's Choco tropical forest, including \textit{Fundación para la Conservación de los Andes Tropicales Reserve and adjacent Reserva Ecológica Mache-Chindul park} (FCAT; 00$^\circ$23'28'' N, 79$^\circ$41'05'' W), \textit{Jama-Coaque Ecological Reserve} (00$^\circ$06'57'' S, 80$^\circ$07'29'' W), \textit{Canande Reserve} (0$^\circ$31'34'' N 79$^\circ$12'47'' W), and \textit{Tesoro Escondido Reserve} (0$^\circ$33'16'' N 79$^\circ$10'31'' W). FCAT is a high diversity humid tropical forest at elevation $\sim$500m, receiving $\sim$3000 mm yr$^{-1}$ precipitation with persistent fog during drier period. Jama-Coaque ranges from the boundary of the tropical moist deciduous/tropical moist evergreen forest at the lower elevations ($\sim$1000 mm precipitation yr$^{-1}$, $\sim$250 m asl) to fog-inundated wet evergreen forests above 580m to 800m. Canande (350–500 m elevation) and Tesoro Escondido ($\sim$200 m elevation) are lowland everwet Choco forests, both

0 papers0 benchmarksImages

Diverse Tools Image Dataset for Machine Learning

This dataset consists of images of various hand and power tools, specifically designed to aid in training and improving AI-based object recognition systems.

0 papers0 benchmarks

Huawei-UK-University-Challenge-Competition-2021 (Huawei UK University Challenge Competition 2021 - TASK2)

<h1>Huawei University Challenge Competition 2021</h1>

0 papers0 benchmarksGraphs

SpaceNet: A Comprehensive Astronomical Dataset

Description:

0 papers0 benchmarks

TURSpider (TURSpider: A Turkish Text-to-SQL Dataset)

TURSpider is a novel Turkish Text-to-SQL dataset that includes complex queries, akin to those in the original Spider dataset. TURSpider dataset comprises two main subsets: a dev set and a training set, aligned with the structure and scale of the popular Spider dataset. The dev set contains 1034 data rows with 1023 unique questions and 584 distinct SQL queries. In the training set, there are 8659 data rows, 8506 unique questions, and corresponding SQL queries.

0 papers0 benchmarks

Tiny Terrors: A Collection of Ticks and Mites

Description:

0 papers0 benchmarks
PreviousPage 675 of 1000Next