TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

BUPTCampus

BUPTCampus is a video-based visible-infrared dataset with approximately pixel-level aligned tracklet pairs and single-camera auxiliary samples.

1 papers0 benchmarks

ITCPR dataset (Image-Text Composed Person Retrieval dataset)

The ITCPR dataset is a comprehensive collection specifically designed for the Zero-Shot Composed Person Retrieval (ZS-CPR) task. It consists of a total of 2,225 annotated triplets, derived from three distinct datasets: Celeb-reID, PRCC, and LAST.

1 papers6 benchmarksImages, Texts

ULS labeled data (UVA laser scanning labelled las data over tropical moist forest classified as leaf or wood points)

UAV Laser Scanning data collected over neotropical forest (Paracou French Guiana). Four flights conducted over one ha plot in 2021 and 2022.

1 papers3 benchmarks3D, LiDAR, Point cloud

First HAREM (Primeiro HAREM)

HAREM, an initiative by Linguateca, boasts a Golden Collection—a meticulously curated repository of annotated Portuguese texts. This resource serves as a pivotal benchmark for evaluating systems in recognizing mentioned entities within documents. It stands as a cornerstone, supporting advancements and innovations in Portuguese language processing research, providing a comprehensive foundation for evaluating system performances and fostering ongoing developments in this domain.

1 papers0 benchmarksTexts

Various URL Datasets (https://github.com/ada-url/url-various-datasets)

Various URL Datasets These are collections of URLs for benchmarking purposes.

1 papers0 benchmarks

NPO (Negative and Positive Obstacles)

The dataset is recorded with an on-vehicle ZED stereo camera in both urban and rural environments

1 papers1 benchmarksRGB-D

ClustMe and ClustML data S1 and S2 (Gaussian Mixture human-labeled data for clustering design and evaluation)

Code and datasets S1 and S2 used in:

1 papers0 benchmarks

Mini HAREM

The MiniHAREM, a reiteration of the 2005 evaluation, used the same methodology and platform. Held from April 3rd to 5th, 2006, it offered participants a 48-hour window to annotate, verify, and submit text collections. Results are available, and the collection used is accessible. Participant lists, submitted outputs, and updated guidelines are provided. Additionally, the HAREM format checker ensures compliance with MiniHAREM directives. Information for the HAREM Meeting, open for registration until June 15th after the Linguateca Summer School in the University of Porto, is also available.

1 papers0 benchmarksTexts

CLCXray (Cutters and Liquid Containers X-ray Dataset)

The CLCXray dataset contains 9,565 X-ray images, in which 4,543 X-ray images (real data) are obtained from the real subway scene and 5,022 X-ray images (simulated data) are scanned from manually designed baggage. There are 12 categories in the CLCXray dataset, including 5 types of cutters and 7 types of liquid containers. Five kinds of cutters include blade, dagger, knife, scissors, swiss army knife. Seven kinds of liquid containers include cans, carton drinks, glass bottle, plastic bottle, vacuum cup, spray cans, tin. The annotations are made in COCO format.

1 papers1 benchmarksImages

EA-HAS-Bench (Energy-Aware Hyperparameter and Architecture Search Benchmark)

We present the first large-scale energy-aware benchmark that allows studying AutoML methods to achieve better trade-offs between performance and search energy consumption, named EA-HAS-Bench. EA-HAS-Bench provides a large-scale architecture/hyperparameter joint search space, covering diversified configurations related to energy consumption. Furthermore, we propose a novel surrogate model specially designed for large joint search space, which proposes a Bezier curve-based model to predict learning curves with unlimited shape and length.

1 papers0 benchmarks

Representative PDE Benchmarks

Given the lack of consensus on a standard setof benchmarks for machine learning of PDEs, we propose a new suite of benchmarks here. Our aims in this regard are to ensure i) sufficient diversity among the types of PDE considered ii) access to training and test data is readily available for rapid prototyping and reproducibility iii) intrinsic computational complexity of problem to make sure that it is worthwhile to design fast surrogates to classical PDE solvers for a particular problem

1 papers0 benchmarks

Algonauts 2023

The Algonauts 2023 Challenge focuses on predicting responses in the human brain as participants perceive complex natural visual scenes. Through collaboration with the Natural Scenes Dataset (NSD) team, the Challenge runs on the largest suitable brain dataset available, opening new venues for data-hungry modeling.

1 papers0 benchmarksImages, MRI

Newsmediabias (Navigating News Narratives: A Media Bias Analysis Dataset)

This is a dataset for news media bias covering different dimensions of the biases: political, hate speech, political, toxicity, sexism, ageism, gender identity, gender discrimination, race/ethnicity, climate change, occupation, spirituality, which makes it a unique contribution

1 papers0 benchmarks

Daily and Sports Activities

The dataset comprises motion sensor data of 19 daily and sports activities each performed by 8 subjects in their own style for 5 minutes. Five Xsens MTx units are used on the torso, arms, and legs.

1 papers0 benchmarksTime series

VGG Physical Property

The VGG Physical Property dataset introduced in the paper is a set of dataset containing positive/negative pairs for understanding of different physical properties, including scene geometry, material, support relation, shadow, occlusion and depth.

1 papers0 benchmarks

Pittsburgh250k

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

MetricInstruct

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

HpVaxFrames

HpVaxFrames includes 64 Vaccine Hesitancy Framings found on Twitter about the HPV vaccines. Language experts annotated tweets as Relevant or Not Relevant, and then further annotated Relevant tweets with Stance towards each framing.

1 papers0 benchmarksTexts

VaccineFrames

Combines CoVaxFrames and HpVaxFrames into a unified dataset of 113 Vaccine Hesitancy Framings found on Twitter about the COVID-19 vaccines and 64 Vaccine Hesitancy Framings found on Twitter about the HPV vaccines. Language experts annotated tweets as Relevant or Not Relevant, and then further annotated Relevant tweets with Stance towards each framing.

1 papers0 benchmarksTexts

Uniswap (Replication Data for: Uniswap Daily Transaction Indices by Network)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksTabular, Time series
PreviousPage 481 of 1000Next