TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

ThermoScenes

Dataset of paired thermal and RGB images comprising ten diverse scenes—six indoor and four outdoor scenes— for 3D scene reconstruction and novel view synthesis (e.g. with NeRF).

1 papers0 benchmarks3D, Hyperspectral images, Images

CFC-DAOD (Caltech Fish Counting – Domain Adaptive Object Detection)

CFC-DAOD is a domain adaptation extension to the Caltech Fish Counting domain generalization benchmark.

1 papers2 benchmarksImages

ACLUE (Ancient Chinese Language Understanding Evaluation)

The Ancient Chinese Language Understanding Evaluation (ACLUE) is an evaluation benchmark focused on ancient Chinese language comprehension. It aims to assess the performance of large-scale language models on understanding ancient Chinese. The benchmark comprises 15 tasks spanning various domains, including lexical, syntactic, semantic, inference, and knowledge. We encourage researchers to use ACLUE to test and enhance their models' abilities in ancient Chinese language understanding. ACLUE's tasks are derived from a combination of manually curated questions from publicly available resources, and automatically generated questions from classical Chinese language corpora. The range of questions span from the Xia dynasty (2070 BCE) to the Ming dynasty (1368 CE). ACLUE adopts a multiple-choice question format for all tasks.

1 papers0 benchmarks

SHREC'20 (Shape Correspondence with Non-Isometric Deformations)

A dataset of highly non-isometric non-rigid quadruped shapes with consensus-based ground-truth correspondences.

1 papers0 benchmarks

APIBench

APIBench is a benchmark dataset designed for evaluating the performance of API recommendation approaches. It was introduced in the paper titled "Revisiting, Benchmarking, and Exploring API Recommendation: How Far Are We?"¹. Let's delve into the details:

1 papers0 benchmarks

MOViD-Amodal

MOViD-A is a video-based synthesized dataset. We create it from MOVi dataset for amodal segmentation. The virtual camera is set to go around the scene, capturing about 24 consecutive frames. We randomly place 10 ∼ 20 static objects that heavily occlude each other in the scene. Finally, we collect 630 and 208 videos for training and testing.

1 papers0 benchmarksVideos

Starlink-on-the-Road Data Set

This data contains throughput, RTT, power consumption and speed data measured with Starlink during mobility setups.

1 papers0 benchmarks

Dataset: Privacy-Preserving Gaze Data Streaming in Immersive Interactive Virtual Reality: Robustness and User Experience.

Collected data from two distinct experiments in immersive, interactive VR where participants performed dynamic tasks as their eye, head, and hand movements were recorded. In the second experiment, a range of real-time privacy mechanisms are applied to eye gaze in real-time.

1 papers0 benchmarksTabular, Tracking

idsprites (Infinite dSprites)

Easily generate simple continual learning benchmarks. Inspired by dSprites.

1 papers0 benchmarksImages

WikiBio GPT-3 Hallucination Dataset

The WikiBio GPT-3 Hallucination Dataset is a benchmark dataset used for hallucination detection. It is based on Wikipedia biographies (WikiBio) and is specifically designed to evaluate the factuality of text generated by large language models like GPT-3¹². Here are some key details about this dataset:

1 papers0 benchmarks

SafeEdit

SafeEdit encompasses 4,050 training, 2,700 validation, and 1,350 test instances. SafeEdit can be utilized across a range of methods, from supervised fine-tuning to reinforcement learning that demands preference data for more secure responses, as well as knowledge editing methods that require a diversity of evaluation texts.

1 papers0 benchmarks

BioDrone

BioDrone is the first bionic drone-based single object tracking benchmark, it features videos captured from a flapping-wing UAV system with a major camera shake due to its aerodynamics. BioDrone highlights the tracking of tiny targets with drastic changes between consecutive frames, providing a new robust vision benchmark for SOT. 1. Large-scale and high-quality benchmark with robust vision challenges 2. Rich challenging factor annotation 3. Videos from Bionic-based UAV 4. Tracking baselines with comprehensive experimental analyses

1 papers0 benchmarksImages, Videos

SOTVerse

SOTVerse is a user-defined task space of single object tracking. It allows users to customize SOT tasks according to their research purposes, which on the one hand makes research more targeted, and on the other hand can significantly improve the efficiency of research.

1 papers0 benchmarksImages, Videos

WikiFactDiff

WikiFactDiff is a dataset designed as a resource to perform atomic factual knowledge updates on language models, with the goal of aligning them with current knowledge. It describes the evolution of factual knowledge between two dates, named T_old and T_new,​ in the form of semantic triples. To enable the possibility of evaluating knowledge algorithms (such as ROME, MEND, MEMIT, etc.), these triples are verbalized and neighbor facts are determined to check for eventual bleedover.

1 papers0 benchmarksTexts

BraTS2019

Multimodal Brain Tumor Segmentation Challenge 2019

1 papers0 benchmarks

DIGITal (Digitally Generated Numerals)

Digitally Generated Numerals (DIGITal) Description The Digitally Generated Numerals (DIGITal) dataset consists of 100,000 image pairs representing digits from 0 to 9. These image pairs include both low and high-quality versions, with a resolution of 128x128 pixels.

1 papers0 benchmarksImages

FABSA (An aspect-based sentiment analysis dataset of Customer Feedback reviews)

FABSA, An aspect-based sentiment analysis dataset in the Customer Feedback space (Trustpilot, Google Play and Apple Store reviews).

1 papers2 benchmarksTexts

Audio de mosquitos Aedes Aegypti (Wing beats)

Dataset Description:

1 papers0 benchmarksAudio

HR-Multiwoz

HR-Multiwoz is a fully-labeled dataset of 550 conversations spanning 10 HR domains to evaluate LLM Agent. It is the first labeled open-sourced conversation dataset in the HR domain for NLP research. Please refer to HR-MultiWOZ: A Task Oriented Dialogue (TOD) Dataset for HR LLM Agent for details about the dataset construction.

1 papers0 benchmarks

HR Extractive Question Answering

Dataset Card HR-Multiwoz is a fully-labeled dataset of 5980 extractive qa spanning 10 HR domains to evaluate LLM Agent. It is the first labeled open-sourced conversation dataset in the HR domain for NLP research. Please refer to HR-MultiWOZ: A Task Oriented Dialogue (TOD) Dataset for HR LLM Agent for details about the dataset construction.

1 papers0 benchmarks
PreviousPage 492 of 1000Next