TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Heteroatom Doped Graphene Supercapacitor

Heteroatom doped graphene supercapacitor feature data is gathered from various literatures for use in machine learning tasks. Main motivation is to optimize supercapacitors and to gain knowledge into models for electrochemistry tasks.

1 papers0 benchmarksTabular

Large Car-following Dataset Based on Lyft level-5: Following Autonomous Vehicles vs. Human-driven Vehicles

Studying how human drivers react differently when following autonomous vehicles (AV) vs. human-driven vehicles (HV) is critical for mixed traffic flow. This dataset contains extracted and enhanced two categories of car-following data, HV-following-AV (H-A) and HV-following-HV (H-H), from the open Lyft level-5 dataset.

1 papers0 benchmarksTime series

Vision-Automatic-Band-Gap-Extractor-Data

Hyperspectral data for a MA/FA lead iodide perovskite gradient.

1 papers0 benchmarks

Vision-Stability-Measurement

Controlled 2-hour degradation optical time series data for a MA/FA lead iodide perovskite gradient.

1 papers0 benchmarks

MID Intrinsics

Intrinsic component extension of MIT Multi-Illumination Dataset proposed in the paper "Intrinsic Image Decomposition via Ordinal Shading", Chris Careaga and Yağız Aksoy, ACM Transactions on Graphics, 2023

1 papers0 benchmarksImages

EyeInfo

The EyeInfo Dataset is an open-source eye-tracking dataset created by Fabricio Batista Narcizo, a research scientist at the IT University of Copenhagen (ITU) and GN Audio A/S (Jabra), Denmark. This dataset was introduced in the paper "High-Accuracy Gaze Estimation for Interpolation-Based Eye-Tracking Methods" (DOI: 10.3390/vision5030041). The dataset contains high-speed monocular eye-tracking data from an off-the-shelf remote eye tracker using active illumination. The data from each user has a text file with data annotations of eye features, environment, viewed targets, and facial features. This dataset follows the principles of the General Data Protection Regulation (GDPR).

1 papers0 benchmarksTabular, Texts, Tracking, Videos

FLOGA (wiLdfire Observations for the Greek Area)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

Concerns and Value Judgments of Stakeholders in the Non-Fungible Tokens (NFTs) Market (Replication Data for: "Centralized or Decentralized?")

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksTabular, Texts, Time series

LPBA40 (LONI Probabilistic Brain Atlas)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks3D, Images, MRI, Medical

Beijing Traffic

The Beijing Traffic Dataset collects traffic speeds at 5-minute granularity for 3126 roadway segments in Beijing between 2022/05/12 and 2022/07/25.

1 papers1 benchmarksTime series

Wikidata5M-SI

Semi-inductive link prediction (LP) in knowledge graphs (KG) is the task of predicting facts for new, previously unseen entities based on context information. Although new entities can be integrated by retraining the model from scratch in principle, such an approach is infeasible for large-scale KGs, where retraining is expensive and new entities may arise frequently. In this paper, we propose and describe a large-scale benchmark to evaluate semi-inductive LP models. The benchmark is based on and extends Wikidata5M: It provides transductive, k-shot, and 0-shot LP tasks, each varying the available information from (i) only KG structure, to (ii) including textual mentions, and (iii) detailed descriptions of the entities. We report on a small study of recent approaches and found that semi-inductive LP performance is far from transductive performance on long-tail entities throughout all experiments. The benchmark provides a test bed for further research into integrating context and textual

1 papers3 benchmarks

FedNLP (FOMC Docs and Speeches)

We collect the various forms of Federal Reserve communications.

1 papers0 benchmarksTexts

formalgeo-imo

IMO-level geometry problem with complete natural language description, geometric shapes, formal language annotations, and theorem sequences annotations.

1 papers0 benchmarks

CARE

https://drive.google.com/file/d/1X_JTfD8Ch-IxmG5VHtKk_xGZT336Fl1Q/view?usp=drive_link

1 papers0 benchmarks

InHARD (Industrial Human Action Recognition Dataset in the Context of Industrial Collaborative Robotics)

We introduce a RGB+S dataset named “Industrial Human Action Recognition Dataset” (InHARD) from a real-world setting for industrial human action recognition with over 2 million frames, collected from 16 distinct subjects. This dataset contains 13 different industrial action classes and over 4800 action samples. The introduction of this dataset should allow us the study and development of various learning techniques for the task of human actions analysis inside industrial environments involving human robot collaborations.

1 papers0 benchmarksRGB-D, Videos

MultiMotionFusion

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

UHGEvalDataset

UHGEvalDataset contains over 5000 news items. It can be used in hallucination evaluation or detection tasks.

1 papers0 benchmarksTexts

BurnMD (A Fire Projection and Mitigation Modeling Dataset)

A dataset composed of 308 medium sized fires from the years 2018-2021, complete with both time series airborne based inference and ground operational estimation of fire extent, and operational mitigation data such as control line construction.

1 papers0 benchmarksEnvironment, Time series

AK_FRAEX - Azure Kinect Frame Extractor demo videos

Video samples recorded in the field using the Azure Kinect DK. These videos are part of the AK-FRAEX software to demonstrate the use of frame extraction tasks. Visit the project site:

1 papers0 benchmarksVideos

cryoPPP (CryoPPP: A Large Expert-Curated Cryo-EM Image Dataset for Machine Learning Protein Particle Picking)

The CryoPPP dataset consists of 34 ground truth data and metadata for 335 EMPIAR IDs. The ground truth data is comprised of a variety of 9893 Micrographs (~300 cryo-EM images per EMPIAR ID) with manually curated ground truth coordinates of picked protein particles. The metadata consists of 1,698,802 high-resolution micrographs deposited in EMPIAR with their respective FPT and Globus data download paths.

1 papers0 benchmarksImages
PreviousPage 480 of 1000Next