TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

MedTurkQuAD: Medical Turkish Question-Answering Dataset

A comprehensive Turkish dataset for question-answering tasks in medical domain

1 papers2 benchmarksTexts

VecCity

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

41598_2022_22531_MOESM2_ESM.xlsx

The datasets used and analysed from the glucose clamp study are available in this Excel file. They include pseudonymised information on the participants, somatometric data, biomarkers of lipid metabolism and parameters of insulin-glucose homeostasis, i.e. concentrations of insulin, glucose and c-peptide as well as data from glucose-clamp experiments, HOMA, SPINA Carb parameters (SPINA-GBeta and SPINA-GR), Matsuda index, insulinogenic index, disposition index and McAuley index.

1 papers0 benchmarksBiomedical, Medical, Tabular, Time series

41598_2022_22531_MOESM1_ESM.dif

The datasets used and analysed from the glucose clamp study are available in this DIF file. They include pseudonymised information on the participants, somatometric data, biomarkers of lipid metabolism and parameters of insulin-glucose homeostasis, i.e. concentrations of insulin, glucose and c-peptide as well as data from glucose-clamp experiments, HOMA, SPINA Carb parameters (SPINA-GBeta and SPINA-GR), Matsuda index, insulinogenic index, disposition index and McAuley index.

1 papers0 benchmarksBiomedical, Medical, Tabular, Time series

Code and Data for Replication

Code and Data for Replication of "Microsimulation Estimates of Decision Uncertainty and Value of Information Are Biased but Consistent"

1 papers0 benchmarksTexts

BPCC-cleaned

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

micro-emotion dataset (Supplementary annotated Goemotions dataset (micro-emotion labels with energy level intensity values ​​(0-10))

The EQN framework is a micro-emotion annotation and detection system that realizes the automatic micro-emotion annotation of text with energy level scores for the first time. The text emotion datasets it annotates are no longer simple single-label or multi-label, but macro-emotions and micro-emotions with continuous values ​​of emotion intensity. The labeling of emotion datasets has changed from discrete to continuous. It plays an important role in the subtle research of emotions in fields such as emotional computing, human-computer alignment, humanoid robots, and psychology.

1 papers0 benchmarks

TemMat (TemMat dataset)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

Experiments data used for evaluating PerfSim simulation accuracy based on sfc-stress workloads (Michel Gokan Khan)

This dataset is being used to evaluate PerfSim accuracy and speed against a real deployment in a Kubernetes cluster based on sfc-stress workloads.

1 papers0 benchmarksTime series

Twitter job title prediction

We introduce a dataset consisting of 1314 samples, including users’ tweets and bios. The user’s job title is found using Wikipedia crawling. The challenge of multiple job titles per user is handled using a semantic word embedding and clustering method. Then, a job prediction method is introduced based on a deep neural network and TF-IDF word embedding. We also use hashtags and emojis in the tweets for job prediction. Results show that the job title of users in Twitter could be well predicted with 54% accuracy in nine categories.

1 papers0 benchmarksTables, Tabular, Texts

ArSen

Sentiment analysis is pivotal in Natural Language Processing for understanding opinions and emotions in text. While advancements in Sentiment analysis for English are notable, Arabic Sentiment Analysis (ASA) lags, despite the growing Arabic online user base. Existing ASA benchmarks are often outdated and lack comprehensive evaluation capabilities for state-of-the-art models. To bridge this gap, we introduce ArSen, a meticulously annotated COVID-19-themed Arabic dataset, and the IFDHN, a novel model incorporating fuzzy logic for enhanced sentiment classification. ArSen provides a contemporary, robust benchmark, and IFDHN achieves state-of-the-art performance on ASA tasks. Comprehensive evaluations demonstrate the efficacy of IFDHN using the ArSen dataset, highlighting future research directions in ASA.

1 papers0 benchmarksTexts

TUD (Table Uniformity Dataset)

Dataset to reproduce results for the paper Detecting CSV file dialects by table uniformity measurement and data type inference, DOI ds240062.

1 papers2 benchmarks

SoliDiffy Differencing Contract Pairs and Edit Scripts

SoliDiffy Differencing Contract Pairs and Edit Scripts Dataset The project creates and maintains two main datasets to assist with research and evaluation of Solidity smart contract differencing:

1 papers0 benchmarksTexts

GPTKB

GPTKB is a large general-domain knowledge base (KB) constructed entirely from a large language model (LLM). It demonstrates the feasibility of large-scale KB construction from LLMs, while highlighting specific challenges arising around entity recognition, entity and property canonicalization, and taxonomy construction.

1 papers0 benchmarksTexts

Relicensing_Forks (Relicensed OSS projects and resulting forks)

The notebooks folder contains basic analysis of the organizational affiliation data for the contributors per open source project.

1 papers0 benchmarks

SIB bioinformatics SPARQL queries

A large collection of human-written natural language questions and their corresponding SPARQL queries over federated bioinformatics knowledge graphs (KGs) collected for several years across different research groups at the SIB Swiss Institute of Bioinformatics. The collection comprises more than 1000 example questions and queries, including 65 federated queries. We propose a methodology to uniformly represent the examples with minimal metadata, based on existing standards. Furthermore, we introduce an extensive set of open-source applications, including query graph visualizations and smart query editors, easily reusable by KG maintainers who adopt the proposed methodology.

1 papers0 benchmarksTexts

IM-SportingBehaviors

IM-SportingBehaviors Dataset Dataset Overview The IM-SportingBehaviors dataset, developed by researchers at Air University Pakistan, provides detailed motion data from participants engaged in various sports activities. This dataset captures human movement through triaxial accelerometers attached to multiple parts of the body, specifically the knee, wrist, and below-neck areas. The dataset includes motion data from six sports: cycling, badminton, skipping, basketball, football, and table tennis, with participants comprising both professional and amateur athletes aged between 20 and 30 years, weighing between 60 and 100 kilograms.

1 papers3 benchmarks

CMACD (Chinese Multi-label Affective Computing Dataset)

This study collected data from the major social media platform Weibo, screening 11,338 valid users from over 50,000 individuals with diverse MBTI personality labels and acquiring 566,900 posts along with the user MBTI personality tags. Using the EQN method, we compiled a multi-label Chinese affective computing dataset that integrates the same user's personality traits with six emotions and micro-emotions, each annotated with intensity levels. This dataset is designed to advance machine recognition of complex human emotions and provide data support for research in psychology, education, marketing, finance, and politics.

1 papers0 benchmarks

BuckTales (A multi-UAV dataset for multi-object tracking and re-identification of wild antelopes)

The first and large scale dataset to solve multi-object tracking and Re-identification problem with wild animals using UAVs.

1 papers0 benchmarks

MovieNet-TeViS

MovieNet-TeViS is a synopsis-storyboard pair dataset to facilitate the Text synopsis to Video Storyboard (TeViS). In the TeViS task, we aim to retrieve an ordered sequence of images from large-scale movie database as video storyboard to visualize an input text synopsis.

1 papers0 benchmarks
PreviousPage 528 of 1000Next