TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

EpiK-Eval (Epistemic Knowledge Evaluation)

Benchmark to evaluate the capability of LMs to consolidate and recall information from multiple training documents.

1 papers0 benchmarksTexts

Genre2Movies (Compositional queries for Movie recommendation)

Genre annotations for movies The file genre2movies.csv contains genre-movie tuples based on Wikidata annotations (https://www.wikidata.org/).

1 papers0 benchmarksGraphs, Ranking, Tabular

A Systematic Review of Open Data in Agriculture

This dataset contains a collection of papers retrieved by using a PRISMA systematic review of Open Data and Public Domain data in Agriculture. This collection of papers uses, creates, or discusses about Open Data and Public Domain.

1 papers0 benchmarks

MVALUE (Multilingual human VALUE dataset)

Multilingual human VALUE(MVALUE) is a multilingual dataset covering 7 concepts of human values: morality, deontology, utilitarianism, fairness, truthfulness, toxicity and harmfulness, each concept subset of it includes positive and negative texts that represent the two opposing directions of the concept. We performed translation on collected human value datasets from English into 15 non-English languages using Google Translate. These languages belong to various language families, including Indo-European (Catalan, French, Indonesian, Portuguese, Spanish), NigerCongo (Chichewa, Swahili), Dravidian (Tamil, Telugu), Uralic (Finnish, Hungarian), Sino-Tibetan (Chinese), Japonic (Japanese), Koreanic (Korean) and Austro-Asiatic (Vietnamese).

1 papers0 benchmarks

IUPAC Standards Online

IUPAC Standards Online is a database built from IUPAC’s standards and recommendations, extracted from the journal Pure and Applied Chemistry (PAC).

1 papers0 benchmarks

FreeMan

FreeMan is the first large-scale multi-view human motion dataset under real scenarios. FreeMan was captured by synchro- nizing 8 smartphones across diverse scenarios. It comprises 11M frames from 8000 sequences, viewed from different perspectives. These sequences cover 40 subjects across 10 different scenarios, each with varying lighting conditions.

1 papers0 benchmarksRGB Video, Videos

MMOS

Mix of Minimal Optimal Sets (MMOS) of dataset has two advantages for two aspects, higher performance and lower construction costs on math reasoning.

1 papers0 benchmarksTexts

ASCAD

ASCAD (ANSSI SCA Database) is a set of databases that aims at providing a benchmarking reference for the SCA community: the purpose is to have something similar to the MNIST database that the Machine Learning community has been using for quite a while now to evaluate classification algorithms performance.

1 papers0 benchmarks

ASCADv2 (ASCAD database version 2)

ASCAD database version 2. This database contained the power consumption of a STM32 Cortex M4 microcrontroller (STM32F303RCT7) during 800.000 random AES encryptions. The AES encryptions are protected with shuffling and affine masking, and the implementation is available on https://github.com/ANSSI-FR/SecAESSTM32. The raw dataset is split into 8 files of 100.000 encryptions, and the extracted dataset contained the 800.000 preprocessed traces with additional metadata.

1 papers0 benchmarks

Intent-based user instruction for electric automation

For the purpose of training and evaluating our intent classification model for electric automation, we curated a dataset consisting of intent-based user instructions. The dataset comprises a total of 14 intents, each associated with approximately 10 user instructions, resulting in a total of 140 instructions for electric automation. The intents were carefully selected to cover a diverse range of control commands and actions commonly encountered in electric automation scenarios. These intents include commands for turning on/off electrical appliances. Each user instruction in the dataset is labeled with its corresponding intent, allowing the model to learn the mapping between input instructions and their intended actions.

1 papers0 benchmarksTexts

ChaosBench

We propose ChaosBench, a large-scale, multi-channel, physics-based benchmark for subseasonal-to-seasonal (S2S) climate prediction. It is framed as a high-dimensional video regression task that consists of 45-year, 60-channel observations for validating physics-based and data-driven models, and training the latter. Physics-based forecasts are generated from 4 national weather agencies with 44-day lead-time and serve as baselines to data-driven forecasts. Our benchmark is one of the first to incorporate physics-based metrics to ensure physically-consistent and explainable models. We establish two tasks: full and sparse dynamics prediction.

1 papers0 benchmarks

Google Local review (Google Local Data)

Description This Dataset contains review information on Google map (ratings, text, images, etc.), business metadata (address, geographical info, descriptions, category information, price, open hours, and MISC info), and links (relative businesses) up to Sep 2021 in the United States.

1 papers0 benchmarksImages, Texts

ViMATH (Vietnamese MATH)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksTexts

ViSR (Vietnamese Synthetic Reasoning)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksTexts

Synthetic Reasoning - Natural Language (Vietnamese Synthetic Reasoning - Natural Language)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksTexts

FFHQH (Flickr-Faces-HQ-Harmonization)

A new dataset for portrait harmonization based on the FFHQ. It contains real images, foreground masks, and synthesized composites.

1 papers2 benchmarksImages

CstomStudio

This dataset is composed of 69 individual objects and 57 meaningful pairs. The objects cover a wide range of categories, including decor item, food, furniture, instrument, jewelry, luggage, person, pet, plant, plushie, scene, thing, toy, transportation, and wearable item.

1 papers0 benchmarks

Ontology Enrichment from Texts (OET): A Biomedical Dataset for Concept Discovery and Placement

A biomedical dataset supporting ontology enrichment from texts, by concept discovery and placement, adapting the MedMentions dataset (PubMed abstracts) with SNOMED CT of versions in 2014 and 2017 under the Diseases (disorder) sub-category and the broader categories of Clinical finding, Procedure, and Pharmaceutical / biologic (CPP) product.

1 papers0 benchmarks

Spurious Boolean Dataset

This is the synthetic dataset that is introduced in the paper https://arxiv.org/abs/2403.03375. It allows fine-grain control on the spurious correlation strength, core and spurious feature hardness/complexity. Parity and staircase function are included in the codebase.

1 papers0 benchmarks

GRD-TRT-BUF-4I Technical Validation Data

This is the static test data from the study "Global Geolocated Realtime Data of Interfleet Urban Transit Bus Iding" collected by GRD-TRT-BUF-4I. test-data-a.csv was collected from December 31, 2023 00:01:30 UTC to January 1, 2024 00:01:30 UTC. test-data-b.csv was collected from January 4, 2024 01:30:30 UTC to January 5, 2024 01:30:30 UTC. test-data-c.csv was collected from January 10, 2024 16:05:30 UTC to January 11, 2024 16:05:30 UTC.

1 papers0 benchmarksTabular
PreviousPage 490 of 1000Next