TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

ReefSet

The full version of ReefSet used in Williams et al. (2024). This dataset contains strongly labeled audio clips from coral reef habitats, taken across 16 unique datasets from 11 countries. This dataset can be used to test transfer learning performance of audio embedding models.

1 papers0 benchmarksAudio

Polarized Film Removal Dataset

The current industrial pipeline includes 315 dynamic industrial scenarios, which can be categorized into three types: QR codes, text, and products. To enhance the diversity, we have different films with diverse material properties, coverage areas, film thicknesses, and levels of wrinkling. The film exhibits significant variability across each scenario. On the other hand, to ensure the stability of the industrial imaging pipeline, we maintained a consistent intensity level for the industrial light source and fixed the distance between the camera and the object flow. This helps to minimize the influence of errors external to the industrial system.

1 papers0 benchmarksImages

Photocast companion dataset

Trained models, evaluation data sets and results, and references to the full data used in the training, validation and testing of the models

1 papers0 benchmarks

PECC (PECC: Problem Extraction and Coding Challenges)

Recent advancements in large language models (LLMs) have showcased their exceptional abilities across various tasks, such as code generation, problem-solving and reasoning. Existing benchmarks evaluate tasks in isolation, yet the extent to which LLMs can understand prose-style tasks, identify the underlying problems, and then generate appropriate code solutions is still unexplored. Addressing this gap, we introduce PECC, a novel benchmark derived from Advent Of Code (AoC) challenges and Project Euler, including 2396 problems. Unlike conventional benchmarks, PECC requires LLMs to interpret narrative-embedded problems, extract requirements, and generate executable code. A key feature of our dataset is the complexity added by natural language prompting in chat-based evaluations, mirroring real-world instruction ambiguities. Results show varying model performance between narrative and neutral problems, with specific challenges in the Euler math-based subset with GPT-3.5-Turbo passing 50% o

1 papers1 benchmarksTexts

Simulated TGC for Muon Angle Fitting

Dataset Details This dataset is primarily created for the work Fast muon tracking with machine learning implemented in FPGA (Arxiv link) that contains ~3M simulated muon events with Geant4. Hits in the muon chamber and ground truth of track angle are saved.

1 papers0 benchmarks

Replication Data for: "DAM" (Replication Data for: "DAM: A Universal Dual Attention Mechanism for Multimodal Timeseries Cryptocurrency Trend Forecasting")

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

BFRD (Bengali Fake Review Dataset)

This is a binary dataset used for Bengali fake review detection in the paper "Bengali Fake Reviews: A Benchmark Dataset and Detection System" accepted in Neurocomputing, a journal published by Elsevier.

1 papers0 benchmarks

IDRCell-100k

We introduce the IDRCell100K image dataset, a collection of biological images, purposefully curated from the extensive and varied Image Data Resource platform. Our selection, based on metadata provided with these experiments, covered various microscopy techniques to encapsulate a diverse array of imaging modalities, ensuring the dataset's breadth in representing biological information. Efforts were made to minimize experimental and imaging biases, striving for a balanced representation up to a feasible extent, thereby reducing dependency on each image modality or experiment.

1 papers0 benchmarksImages

RMOT-223

In this dataset, various objects are arranged on a white table. A UR5e robot picks and place a target object specified on the title of the video/image sequence. Videos under auto- folder are collected with automatic operation of the robot. Videos under human- folders are collected with the tele-operation of the robot. Ground-truth tracking bounding boxes are generated with STARK, and when the target exits the camera frame, the bounding box estimation is switched to [-1, -1, -1, -1], indicating target not shown.

1 papers0 benchmarksImages, Videos

Computer Codes in Matlab and Fortran

The computer codes (in Matlab or Fortran) can be downloaded from the website: people.maths.ox.ac.uk/erban/Education/

1 papers0 benchmarks

ViTHSD (Vietnamese Targeted-Hate-Speech-Detection)

A Vietnamese dataset for hate speech detection by the specific target. The dataset contains 10,000 comments, each comment has 05 targets with three relevant hateful levels.

1 papers0 benchmarksTexts

Phase 1 dataset

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksImages

Phase 2 dataset

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksImages

Labels

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksTexts

MQuAKE-2002

A new knowledge editing dataset removing the incorrect annotation mistakes in the prior dataset MQuAKE.

1 papers0 benchmarks

MQuAKE-hard

A new knowledge editing benchmark including the challenging instances collected from MQuAKE.

1 papers0 benchmarks

PerezGaldos (Single Spanish Speaker Dataset)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksAudio

DivaTrack dataset

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

NES-VMDB

NES-VMDB is a dataset containing 98,940 gameplay videos from 389 NES games, each paired with its original soundtrack in symbolic format (MIDI). NES-VMDB is built upon the Nintendo Entertainment System Music Database (NES-MDB), encompassing 5,278 music pieces from 397 NES games.

1 papers0 benchmarksMidi, Videos

IDD-X

Intelligent vehicle systems require a deep understanding of the interplay between road conditions, surrounding entities, and the ego vehicle's driving behavior for explainable driving decision-making and safe and efficient navigation. This is particularly critical in developing countries where traffic situations are often dense and unstructured with heterogeneous road occupants. Existing datasets, predominantly geared towards structured and sparse traffic scenarios, fall short of capturing the complexity of driving in such environments. To fill this gap, we present IDD-X, a large-scale dual-view driving video dataset. With 697K bounding boxes, 9K important object tracks, and 1-12 objects per video, IDD-X offers comprehensive ego-relative annotations for multiple important road objects covering 10 categories and 19 explanation label categories. The dataset also incorporates rearview information to provide a more complete representation of the driving environment. We also introduce custo

1 papers0 benchmarksTexts, Videos
PreviousPage 499 of 1000Next