TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

LLM Generated Spear Phishing Emails

This dataset comprises high-quality, targeted spear-phishing emails created using a proprietary system that harnesses the power of LLMs and knowledge graphs. The primary purpose of releasing this dataset is to promote and facilitate further research in the field of spear-phishing detection.

1 papers0 benchmarksTexts

Experimental materials

Contains materials used in the experiments, raw results, and analysis scripts.

1 papers0 benchmarks

Household Waste Dataset

Collection of images of garbages grouped into 10 classes (metal, glass, biological, paper, battery, trash, cardboard, shoes, clothes, and plastic). The number of files in respective classes is as follows:

1 papers0 benchmarks

LoWRA Bench Dataset

The LoRA Weight Recovery Attack (LoWRA) Bench is a comprehensive benchmark designed to evaluate Pre-Fine-Tuning (Pre-FT) weight recovery methods as presented in the "Recovering the Pre-Fine-Tuning Weights of Generative Models" paper.

1 papers0 benchmarks

Deep Deep Learning With BART (Trained Weights and Example Data)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksImages, MRI, Medical

ScreenQA Short

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

Multilingual Fever

https://github.com/lingo-iitgn/XME

1 papers0 benchmarks

Analyzing Reward Dynamics and Decentralization in Ethereum 2.0 (Replication Data for: Analyzing Reward Dynamics and Decentralization in Ethereum 2.0)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

Arxiv-5Lbench

Arxiv dataset

1 papers0 benchmarks

LeQua2022 (Learning to Quantify Dataset 2024)

This is the dataset used in the 1st data challenge on Learning to Quantify. It is designed for the comparative evaluation of methods for “learning to quantify” in textual datasets, i.e., methods for training predictors of the relative frequencies of the classes of interest in sets of unlabelled textual documents. These predictors (called “quantifiers”) are required to issue predictions for several such sets, some of them characterized by class frequencies radically different from the ones of the training set.

1 papers0 benchmarks

BiMed1.3M

The dataset covers three types of medical interactions in both English and Arabic: - Multiple-choice question answering (MCQA), focusing on specialized medical knowledge. - Open question answering (QA), including real-world consumer questions. - MCQA-Grounded multi-turn chat conversations for dynamic exchanges.

1 papers0 benchmarks

MVME (Multi-View Medical Evaluation Benchmark)

The benchmark assesses the real-time interactive consultation capabilities of LLMs across three critical dimensions. We collect Chinese medical records across diverse departments online.

1 papers0 benchmarksTexts

NaturalTransform (Natural Semantic-preserving Transformations)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

V-Rank (SIMULATED AIRCRAFT TRAJECTORY FOR THEORETICAL VELOCITY RANKING)

Abstract This data set is a data set used for aircraft theoretical velocity ranking. Four sensors are randomly arranged in a 1*1 square map, and three aircraft will fly over the map coverage area at the same time. The velocity of the aircraft is simulated by a random process. The theoretical velocities of the three aircraft are similar, and the velocity of the aircraft will be disturbed during actual flight, causing large fluctuations, so that it is difficult to distinguish the theoretical velocity order of the aircraft flying into the map. The coverage area of the sensor is circular with a fixed radius. The four sensors have a unified detection interval event and will detect the position of the aircraft within the coverage area with unified accuracy. The target task is to reason the theoretical velocity ranking of three aircraft through the trajectory data collected by the sensors.

1 papers0 benchmarksTracking

BIDS CHB-MIT Scalp EEG Database

This dataset is a BIDS-compatible version of the CHB-MIT Scalp EEG Database. It reorganizes the file structure to comply with the BIDS specification. To this effect:

1 papers0 benchmarksEEG, Medical, Time series

BIDS Siena Scalp EEG Database

This dataset is a BIDS compatible version of the Siena Scalp EEG Database. It reorganizes the file structure to comply with the BIDS specification. To this effect:

1 papers0 benchmarksEEG, Medical, Time series

Siena Scalp EEG Database (Physionet Siena Scalp EEG Database)

The database consists of EEG recordings of 14 patients acquired at the Unit of Neurology and Neurophysiology of the University of Siena. Subjects include 9 males (ages 25-71) and 5 females (ages 20-58). Subjects were monitored with a Video-EEG with a sampling rate of 512 Hz, with electrodes arranged on the basis of the international 10-20 System. Most of the recordings also contain 1 or 2 EKG signals. The diagnosis of epilepsy and the classification of seizures according to the criteria of the International League Against Epilepsy were performed by an expert clinician after a careful review of the clinical and electrophysiological data of each patient.

1 papers0 benchmarksEEG, Medical, Time series

SeizeIT1

This dataset is obtained during an ICON project (2017-2018) in collaboration with KU Leuven (ESAT-STADIUS), UZ Leuven, UCB, Byteflies and Pilipili. The goal of this project was to design a system using Behind the ear (bhE) EEG electrodes for monitoring the patient in a home environment. This way, a nice balance can be found between sufficient accuracy of seizure detection algorithms (because EEG is used) and wearability (bhe EEG is relatively subtle, similar to a hear-aid device). The dataset acquired in the hospital during presurgical evaluation. During such presurgical evaluation, neurologists try to see if a specific part of the brain is causing the seizures, and if so, if that part of the brain can be removed during surgery. During the presurgical evaluation, patients are monitored using the vEEG for multiple days (typically a week). Patients are however restricted to move within their room because of the wiring and video analysis. In this dataset, following data is available per p

1 papers0 benchmarksEEG, Medical, Time series

CausalGym

SyntaxGym, adapted for interventional interpretability.

1 papers1 benchmarksTexts

HatefulDiscussions

Multi-Modal Hate Speech Detection with Graph Context.

1 papers0 benchmarksGraphs, Images, Texts
PreviousPage 488 of 1000Next