TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

SCOPE Dataset

A Chinese sign language dataset that includes dialogue information.

1 papers0 benchmarksTexts, Videos

HSIRS (High-quality Spectral Image Resonstruction and Segmentation Dataset)

We introduce HSIRS, a large scale dataset of hyper-spectral images along with corresponding manually annotated segmentation maps for material characterization and classification based on spectral signature. Such data can be used to simulate any type of spectrometer and to train DNNs end-to-end for spectral reconstruction and image segmentation tasks. HSIRS features scenes containing real and fake (made of polyester, plastic or ceramic) food items with different backgrounds and scene layouts, some scenes contain also color checkers. Spectral bands are sequentially captured using a VariSpecTM tunable color filter and the scene is illuminated with 4 Halogen light sources.

1 papers0 benchmarksHyperspectral images, Images

Open RAN Commercial Traffic Twinning Dataset

Dataset of cross-layer Radio Access Network (RAN) Key Performance Measurements (KPMs) and protocol stack logs collected on an Open RAN deployment instantiated on Colosseum with traffic twinned from that of commercial cellular traces. The dataset includes Base Station (BS)- and User Equipment (UE)-level KPMs from PHY, MAC, and App layers under different RAN configurations representative of AI/ML control policies, number of UEs, and traffic demand. The fine-grained metrics of the dataset make it possible to understand the connection between PHY and MAC KPMs measured at the BS and UEs, control policies, and end-to-end and App-layer KPMs that reflect user experience.

1 papers0 benchmarksTime series

Guitar Playing Motion Dataset

Diverse guitar-playing motions about 1 hour long, including: • 12 major scales, • chromatic scales, • diverse chords, • arpeggios, • strumming and picking, • bends, • sliding, • vibrato, • palm mute, • natural harmonics, • artificial harmonics, • hammer-ons and pull-offs.

1 papers0 benchmarks

GTA-UAV

GTA-UAV dataset provides a large continuous area dataset (covering 81.3km<sup>2</sup>) for UAV visual geo-localization, expanding the previously aligned drone-satellite pairs to arbitrary drone-satellite pairs to better align with real-world application scenarios. Our dataset contains:

1 papers0 benchmarksImages

CoyPu-Mini-Text2Sparql

This Dataset contains pairs off textual natural language questions and SPARQL queries on a small subset of the CoyPu KnowledgeGraph (https://coypu.org/ergebnisse/knowledge-graph)

1 papers0 benchmarks

Organizational-Text2Sparql

This Dataset contains pairs off textual natural language questions and SPARQL queries on a small organizational graph(https://github.com/AKSW/AI-Tomorrow-2023-KG-ChatGPT-Experiments/blob/main/FoafVcardOrg/foaf-vcard-org-data.ttl) which was introduced in "LLM-assisted knowledge graph engineering: Experiments with chatgpt" by L.-P. Meyer et al. 2023 (DOI 10.1007/978-3-658-43705-3_8)

1 papers0 benchmarks

RRG (Russian RST dataset from GUM v9.1 corpus)

Parallel version of annotations in GUM RST v9.1.

1 papers0 benchmarksTexts

RRT (RuRSTreebank)

RST corpus for Russian.

1 papers0 benchmarksTexts

diderot_1751_wd

Diderot’s Encyclopédie is a reference work from XVIIIth century in Europe that aimed at collecting the knowledge of its era. This repository hosts an annotated dataset of more than 10,400 of the Encyclopédie entries with Wikidata identifiers enabling us to connect these entries to the Wikidata graph. The dataset can serve to train and evaluate named entity solvers.

1 papers0 benchmarksTexts

Orchid2024

Orchid2024 is a fine-grained classification dataset specifically designed for Chinese Cymbidium orchid cultivars. It includes data collected from 20 cities across 12 provincial administrative regions in China and encompasses 1,269 cultivars from 8 Chinese Cymbidium orchid species and 6 additional categories, totaling 156,630 images. The dataset covers nearly all common Chinese Cymbidium cultivars currently found in China, with its fine granularity and focus on the real world making it a unique and practical resource for researchers and practitioners.

1 papers0 benchmarksImages

RealArt-6

This is the official dataset collected for to test the sim-to-real transfer. It contains 6 articulated object instances, each captured from 20 camera views under 5 states in scenarios with and without background, as well as presence or absence of distractors.

1 papers0 benchmarks3D, Point cloud

wonderbread

A benchmark + dataset for evaluating multimodal models on business process management (BPM) tasks.

1 papers0 benchmarks

FUSU

FUSU dataset covers 5 whole urban areas, 847 km^2 located in the north and south of China, with 17 land use and land cover (LULC) classes and over 170K images and 30 billion pixels of annotations, supporting segmentation, change detection and domain adaptation tasks. This data comprises 2 parts:

1 papers0 benchmarks

Aria Everyday Objects

A small-scale, real-world Project Aria dataset with high quality static 3D oriented bounding boxs annotations.

1 papers6 benchmarks3D, Point cloud, Videos

ARLBench Data

Around 90k different RL runs: 256 hyperparameter configurations for 10 seeds each across a total of 3 algorithms (PPO, SAC, DQN) and 22 environments.

1 papers0 benchmarks

GLARE (Guided LexRank for Advanced Retrieval in Legal Analysis)

The Guided Lexrank algorithm is applied to dataset special_appeal.csv to summarize the texts of legal documents. The obtained summary and the texts of the topics contained in dataset themes.csv are submitted to the BM25 algorithm for similarity assessment. From a list of topics, the GLARE method produces a ranking with suggested topics for a given document.

1 papers0 benchmarksTexts

MemBench

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksImages

ALLO (Anomaly Localization in Lunar Orbit)

ALLO is an anomaly detection and localization dataset for space stations in lunar orbit. Synthetically rendered using Blender, ALLO provides realistic images of what a robotic manipulator on a space station will encounter including possible anomalies.

1 papers0 benchmarksImages

Group LAW (LAW | SuiteSparse Matrix Collection)

URL: https://sparse.tamu.edu/LAW

1 papers0 benchmarks
PreviousPage 520 of 1000Next