TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

NSC (National Speech Corpus)

The National Speech Corpus (NSC) is a significant initiative led by the Info-communications and Media Development Authority (IMDA) of Singapore. It serves as the first large-scale Singapore English corpus and aims to be a valuable resource for researchers and developers working on automatic speech recognition (ASR) technology and other speech-related applications¹²³.

0 papers0 benchmarks

Safety Prompts

The Safety Prompts dataset is a valuable resource for evaluating and enhancing the safety of large language models (LLMs) in the Chinese language. It consists of carefully crafted prompts that align model outputs with human values, specifically focusing on safety assessment.

0 papers0 benchmarks

SPaRCNet

Seizures and seizure-like rhythmic and periodic brain activity known as “ictal-interictal-injury continuum” (IIIC) patterns are frequently detected during brain monitoring with electroencephalography (EEG) in patients with epilepsy or critical illness. Prior efforts to automate detection of IIIC patterns have been limited by lack of large well-annotated datasets to train/evaluate algorithms, and there have been only a few attempts to detect IIIC events other than seizures. The IIIC dataset includes 50,697 labeled EEG samples from 2,711 patients’ and 6,095 EEGs that were annotated by physician experts from 18 institutions. These samples were used to train SPaRCNet (Seizures, Periodic and Rhythmic Continuum patterns Deep Neural Network), a computer program that classifies IIIC events with an accuracy matching clinical experts.

0 papers0 benchmarks

UAV-C

UAV tracking dataset with common corruptions tailored for unmannered aerial video

0 papers0 benchmarks

BIGOS (Benchmark Intended Grouping of Open Speech)

The Benchmark Intended Grouping of Open Speech (BIGOS) is a novel corpus specifically designed for Polish Automatic Speech Recognition (ASR) systems. This initial version of the benchmark comprises 1,900 audio recordings from 71 distinct speakers, sourced from 10 publicly available speech corpora¹²³.

0 papers0 benchmarks

Z-Bench

Z-Bench is a fascinating Chinese language model prompt dataset developed by an enthusiastic AI-focused team at Zhenfund. Let me share some intriguing details about it:

0 papers0 benchmarks

Indonesian Twitter Emotion Dataset

This dataset contains 4,403 Indonesian tweets that have been labeled into five emotion classes: love, anger, sadness, joy, and fear. Each line in the dataset consists of a tweet and its respective emotion label separated by a semicolon (,). The first line serves as a header. Pre-processing has been applied to the tweets, including replacing usernames (@) with [USERNAME], URLs/hyperlinks (http://… or https://…) with [URL], and sensitive numbers (e.g., phone numbers, invoice numbers) with [SENSITIVE-NO].

0 papers0 benchmarks

Supplementary software and data: Bayesian multi-exposure image fusion for robust high dynamic range ptychography

Accompanying software package and data for the publication titled "Bayesian multi-exposure image fusion for robust high dynamic range ptychography".

0 papers0 benchmarks

THAR Dataset (Targeted Hate Speech Against Religion)

The increase in religiously motivated hate on social media is clear and ongoing. These platforms have become fertile ground for the dissemination of hate speech directed at religious communities, resulting in tangible repercussions in the real world. Much of the current research concerning the automated identification of hateful content on social media focuses on English-language content. There is comparatively less exploration in low-resource languages such as Hindi. As social media users increasingly utilize their regional languages for expression, it becomes crucial to dedicate appropriate research efforts to hate speech detection in these languages.

0 papers0 benchmarksTexts

Mudestreda (Mudestreda Multimodal Device State Recognition Dataset)

Mudestreda Multimodal Device State Recognition Dataset obtained from real industrial milling device with Time Series and Image Data for Classification, Regression, Anomaly Detection, Remaining Useful Life (RUL) estimation, Signal Drift measurement, Zero Shot Flank Took Wear, and Feature Engineering purposes.

0 papers0 benchmarksAudio, Images, Time series

Google Research Football

Google Research Footbal is a new reinforcement learning environment where agents are trained to play football in an advanced, physics-based 3D simulator. The resulting environment is challenging, easy to use and customize, and it is available under a permissive open-source license. In addition, it provides support for multiplayer and multi-agent experiments.

0 papers0 benchmarks

academy 3 vs 1 with keeper on GRF (academy 3 vs 1 with keeper on Google Research Football)

academy 3 vs 1 with keeper on Google Research Football

0 papers0 benchmarks

YouTube Subtitles

YT_subtitles is a remarkable tool designed for building a dataset from YouTube subtitles. Let me break it down for you:

0 papers0 benchmarks

Hacker News

The Hacker News dataset provides a valuable glimpse into the tech industry's landscape. It encompasses posts from Y Combinator's social news website, spanning from 2006 to late 2017¹. Here are some key points about this dataset:

0 papers0 benchmarks

NIH Grant Abstracts (ExPORTER)

NIH Grant Abstracts: ExPORTER is a valuable resource for researchers and data enthusiasts. It serves as an open data repository containing administrative information about NIH-funded research projects. Here are some key points about ExPORTER:

0 papers0 benchmarks

Nordjylland News

The Nordjylland News dataset is a collection of news articles from Northern Jutland in Denmark. It provides valuable information about events, developments, and stories specific to that region. Here are some key details about this dataset:

0 papers0 benchmarks

no sammendrag

The "no-sammendrag" dataset is a collection of summarized texts in Norwegian. It covers various topics and is useful for tasks related to summarization. The dataset contains documents along with their corresponding summaries. Let me provide more details about it:

0 papers0 benchmarks

Dutch Social

The Dutch Social Dataset is a collection of tweets primarily written in Dutch. Let me provide you with some details about this dataset:

0 papers0 benchmarks

Nordjylland News Summarization

The Nordjylland News summarization dataset is a collection of news articles from the Nordjylland region in Denmark, along with their corresponding summaries. This dataset is valuable for training and evaluating text summarization models. Let's delve into the details:

0 papers0 benchmarks

Absabank-Imm

The Swedish ABSAbank-Imm 1.1 is an annotated corpus designed for aspect-based sentiment analysis related to immigration in Sweden. Let's delve into the details:

0 papers0 benchmarks
PreviousPage 658 of 1000Next