TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

AI-generated Twitter Timelines

This project contains instructions and codes to reconstruct a dataset for the development and evaluation of forensic tools for detecting machine-generated text in social media.

1 papers0 benchmarksTexts

PPTC (PowerPoint Task Completion (PPTC))

Recent evaluations of Large Language Models (LLMs) have centered around testing their zero-shot/few-shot capabilities for basic natural language tasks and their ability to translate instructions into tool APIs. However, the evaluation of LLMs utilizing complex tools to finish multi-turn, multi-modal instructions in a complex multi-modal environment has not been investigated. To address this gap, we introduce the PowerPoint Task Completion (PPTC) benchmark to assess LLMs' ability to create and edit PPT files based on user instructions. It contains 279 multi-turn sessions covering diverse topics and hundreds of instructions involving multi-modal operations.

1 papers0 benchmarks

Unified NSL Data Feb 2024 (Unified NSL Data until Feb 2024)

The dataset is hosted on GitHub: https://github.com/ucsdsysnet/nsl-empirical-analysis and explained in our paper.

1 papers0 benchmarks

basqueparl (BasqueParl)

This repository contains BasqueParl, a bilingual corpus for political discourse analysis. It covers transcriptions from the Parliament of the Basque Autonomous Community for eight years and two legislative terms (2012-2020), and its main characteristic is the presence of Basque-Spanish code-switching speeches.

1 papers0 benchmarks

AlgoPuzzleVQA

We introduce the novel task of multimodal puzzle solving, framed within the context of visual question-answering. We present a new dataset, AlgoPuzzleVQA designed to challenge and evaluate the capabilities of multimodal language models in solving algorithmic puzzles that necessitate both visual understanding, language understanding, and complex algorithmic reasoning. We create the puzzles to encompass a diverse array of mathematical and algorithmic topics such as boolean logic, combinatorics, graph theory, optimization, search, etc., aiming to evaluate the gap between visual data interpretation and algorithmic problem-solving skills. The dataset is generated automatically from code authored by humans. All our puzzles have exact solutions that can be found from the algorithm without tedious human calculations. It ensures that our dataset can be scaled up arbitrarily in terms of reasoning complexity and dataset size. Our investigation reveals that large language models (LLMs) such as GPT

1 papers1 benchmarksImages, Texts

HistGen WSI-Report Dataset

This dataset is composed of 7,753 pairs of whole slide images and their corresponding diagnostic reports, extracted from the TCGA platform and refined with large language models. This dataset aims to boost the field of automated histopathology report generation by providing a new publicly available evaluation benchmark. See HistGen paper (see https://arxiv.org/pdf/2403.05396.pdf for reference) for a more detailed description of this dataset.

1 papers1 benchmarksImages, Texts

Battery surface temperature

battery surface temperature from 16 sensors.

1 papers0 benchmarks

ATD-Dataset (Auto-Tune Detection Dataset (ATD-Dataset))

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksAudio

SECO (Seasonal Contrast)

Read more about the dataset here: https://github.com/ServiceNow/seasonal-contrast

1 papers0 benchmarks

Omnicount-191

To effectively evaluate OmniCount across open-vocabulary, supervised, and few-shot counting tasks, a dataset catering to a broad spectrum of visual categories and instances featuring various visual categories with multiple instances and classes per image is essential. The current datasets, primarily designed for object counting focusing on singular object categories like humans and vehicles, fall short for multi-label object counting tasks. Despite the presence of multi-class datasets like MS COCO, their utility is limited for counting due to the sparse nature of object appearance. Addressing this gap, we created a new dataset with 30,230 images spanning 191 diverse categories, including kitchen utensils, office supplies, vehicles, and animals. This dataset, featuring a wide range of object instance counts per image ranging from 1 to 160 and an average count of 10, bridges the existing void and establishes a benchmark for assessing counting models in varied scenarios.

1 papers1 benchmarksImages, Texts

Dark Corner Artifact Masks for ISIC Images

A set of over 4000 dark corner artifact / vignette masks that can be applied to ISIC skin lesion images to test for the effect of such artifacts in classification tasks.

1 papers0 benchmarks

MPII Human Pose Descriptions

The MPII Human Pose Descriptions dataset extends the widely-used MPII Human Pose Dataset with rich textual annotations. These annotations are generated by various state-of-the-art language models (LLMs) and include detailed descriptions of the activities being performed, the count of people present, and their specific poses.

1 papers0 benchmarksTexts

Qubit energy relaxation versus frequency and time

Qubit energy relaxation versus frequency and time

1 papers0 benchmarks

VIRDO Dataset (VIRDO Simulated Kitchen Utensil Deformation Dataset)

From https://github.com/MMintLab/VIRDO/blob/master/data/dataset_readme.txt,

1 papers0 benchmarks3D, Point cloud

HH Red Teaming

The HH Red Teaming dataset comprises two distinct types of data, each serving a unique purpose:

1 papers0 benchmarks

MemoTrap

The MemoTrap dataset is a diagnostic dataset designed to assess whether language models (LMs) fall into memorization traps. These traps occur when LMs memorize specific examples from their training data instead of generalizing effectively. The dataset aims to probe the susceptibility of LMs to this phenomenon.

1 papers0 benchmarks

MNumGLUESub

The MNumGLUESub dataset is a multi-task arithmetic reasoning benchmark. It is part of the NumGLUE collection, which focuses on evaluating mathematical reasoning abilities in various languages. Let's break down the details:

1 papers0 benchmarks

WFO (World Flora Online)

The World Flora Online (WFO) is a comprehensive initiative that aims to create an online flora of all known plants. Here are some key details about the WFO:

1 papers0 benchmarks

Truss dataset

Data used in the work ”Unifying the design space and optimizing linear and nonlinear truss metamaterials by generative modeling”. This dataset contains 965,736 truss structures and their corresponding effective homogenized stiffness.

1 papers0 benchmarks

Integrated Drone Dataset for Semantic Segmentation (IDD)

Semantic segmentation of drone images is critical for various aerial vision tasks as it provides essential semantic details to understand scenes on the ground. Ensuring high accuracy of semantic segmentation models for drones requires access to diverse, large-scale, and high-resolution datasets, which are often scarce in the field of aerial image processing. While existing datasets typically focus on urban scenes and are relatively small, our Varied Drone Dataset (VDD) addresses these limitations by offering a large-scale, densely labeled collection of 400 high-resolution images spanning 7 classes. This dataset features various scenes in urban, industrial, rural, and natural areas, captured from different camera angles and under diverse lighting conditions. We also make new annotations to UDD and UAVid, integrating them under VDD annotation standards, to create the Integrated Drone Dataset (IDD). It's expected that our dataset will generate considerable interest in drone image segmenta

1 papers0 benchmarks
PreviousPage 491 of 1000Next