TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Angry Tweets

The Angry Tweets dataset is a collection of anonymized Danish Twitter data that has been annotated for sentiment analysis through crowd-sourcing. Here are some key details about the dataset:

1 papers0 benchmarks

10kGNAD (Ten Thousand German News Articles Dataset)

The 10kGNAD dataset is intended to solve part of this problem as the first German topic classification dataset. It consists of 10273 German-language news articles from an Austrian online newspaper categorized into nine topics. These articles are a now unused part of the One Million Posts Corpus.

1 papers0 benchmarks

MusicCrowd

A Crowdsourced Multi-Domain Music Dataset of Europeana Collections

1 papers0 benchmarks

Touché23-ValueEval

The Touché23-ValueEval Dataset is a collection of arguments used for identifying human values behind those arguments. It was created by collecting 9324 arguments from 6 diverse sources, including religious texts, political discussions, free-text arguments, newspaper editorials, and online democracy platforms. Each argument was annotated by 3 crowdworkers for 54 values.

1 papers0 benchmarks

cCOVID-News

The cCOVID-News dataset is a publicly available Chinese text retrieval dataset created from COVID-19 news articles. It contains a collection of text data related to COVID-19, and it is used as part of the out-of-domain evaluation for the DuReader retrieval benchmark.

1 papers0 benchmarks

WMT 2021 - Multilingual Low-Resource Translation for Indo-European Languages

The Multilingual Low-Resource Translation task for Indo-European Languages, part of the EMNLP 2021 Conference, focused on improving machine translation in the cultural heritage domain for North-Germanic and Romance languages. It aimed to explore data transferability across related languages, prioritizing low-resource languages while allowing training in high-resource languages. The task had two subtasks: translating Europeana thesis abstracts and descriptions for North-Germanic languages, and translating Wikipedia cultural heritage articles for Romance languages. The task encouraged using diverse data sources and provided additional resources like lexicons and validation sets. Evaluation was based on translation quality, emphasizing multilinguality and resource efficiency in machine translation.

1 papers0 benchmarks

TACO-BAAI (Topics in Algorithmic Code generation dataset)

TACO (Topics in Algorithmic Code generation dataset) is a dataset focused on algorithmic code generation, designed to provide a more challenging training dataset and evaluation benchmark for the code generation model field. The dataset consists of programming competition problems that are more difficult and closer to real programming scenarios. It emphasizes improving or evaluating the model's understanding and reasoning abilities in practical application scenarios, rather than just implementing predefined function functionalities.

1 papers1 benchmarksTexts

TinyShakespeare

The Tiny Shakespeare corpus is a dataset that contains 40,000 lines of Shakespeare from a variety of his plays. The Tiny Shakespeare corpus is a popular choice for training language models due to its manageable size and the complexity of Shakespeare's language. It provides a good balance between computational efficiency and the ability to generate interesting text.

1 papers0 benchmarks

PyranometerDataAtUTEQ2020_2022 (Pyranometer Data At UTEQ 2020-2022)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

TuPyE-Dataset (Portuguese Hate Speech Expanded Dataset)

TuPyE, an enhanced iteration of TuPy, encompasses a compilation of 43,668 meticulously annotated documents specifically selected for the purpose of hate speech detection within diverse social network contexts. This augmented dataset integrates supplementary annotations and amalgamates with datasets sourced from Fortuna et al. (2019), Leite et al. (2020), and Vargas et al. (2022), complemented by an infusion of 10,000 original documents from the TuPy-Dataset.

1 papers0 benchmarksTexts

KAgentBench

KAgentBench is a benchmark dataset of over 3,000 human-edited, automated evaluation data for testing agent capabilities, with evaluation dimensions including planning, tool use, reflection, concluding, and profiling.

1 papers0 benchmarks

Multi-species DNA

Latent DNA Diffusion Dataset

1 papers0 benchmarks

Sequential Instructions

This is the sequential instructions dataset from Understanding the Effects of RLHF on LLM Generalisation and Diversity. The dataset is in the alpaca_eval format.

1 papers0 benchmarksTexts

ODSI-DB (ODSI-DB – Oral and Dental Spectral Image Database)

ODSI-DB is an image database of oral and dental reflectance spectral images of human test subjects. Image sets of the test subjects contain the front-view and the occlusal surfaces of lower and upper teeth, oral mucosa, and face surrounding the mouth. Other features-of-interest have been imaged on case-by-case basis. The spectral images in the database have been annotated by dental experts.

1 papers0 benchmarksImages

January 2, 2024 (v1) Software Open Software and DataSet of "A QA-SQP assisted FE for non-linear and history-dependent mechanics"

Click to #Development of QA-SQP for non-linear and history-dependent mechanical problems

1 papers0 benchmarks

Machine_Mindset_MBTI_dataset

Dataset introduction There are four dimension in MBTI. And there are two opposite attributes within each dimension.

1 papers0 benchmarksTexts

IoT ENVIRONMENT DATASET

ABSTRACT Recently, the technology of the fourth revolution has given the characteristics of things constantly expanding, and everything, including people, things, people, and the environment, is connected based on the Internet. In particular, the network structure is connected to various IoT devices and is changing from wired to wireless. Unlike users who operated each device, other devices can now be operated through gateways inside and outside the smart home. However, these changes have created an environment vulnerable to external attacks, and when an attacker accesses a gateway, he can attempt various attacks, including Port scans, OS&Service detection, and DoS attacks on IoT devices. Therefore, we disclose the dataset below to promote security research on IoT.

1 papers0 benchmarks

Forex News Annotated Dataset for Sentiment Analysis

This dataset contains news headlines relevant to key forex pairs: AUDUSD, EURCHF, EURUSD, GBPUSD, and USDJPY. The data was extracted from reputable platforms Forex Live and FXstreet over a period of 86 days, from January to May 2023. The dataset comprises 2,291 unique news headlines. Each headline includes an associated forex pair, timestamp, source, author, URL, and the corresponding article text. Data was collected using web scraping techniques executed via a custom service on a virtual machine. This service periodically retrieves the latest news for a specified forex pair (ticker) from each platform, parsing all available information. The collected data is then processed to extract details such as the article's timestamp, author, and URL. The URL is further used to retrieve the full text of each article. This data acquisition process repeats approximately every 15 minutes.

1 papers0 benchmarksTexts

UCCS (UnConstrained College Students Dataset)

Unconstrained Face Detection and Open-Set Face Recognition Challenge

1 papers0 benchmarks

Polarimetric Imaging for Perception

The dataset includes polarimetric, RGB and depth automotive (on the road) data.

1 papers0 benchmarksImages, LiDAR
PreviousPage 484 of 1000Next