TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Hawk Annotation Dataset

Hawk Annotation Dataset includes language descriptions specifically for anomaly scenes in seven existing video anomaly datasets. These seven datasets include a variety of anomalous scenarios such as crime (UCF-Cirme), campus (ShanghaiTech and CUHK Avenue), pedestrian walkways (UCSD Ped1 and Ped2), traffic (DoTA), and human behavior (UBnormal). With the support of these visual scenarios, this dataset can perform comprehensive fine-tuning for various abnormal scenarios, being closer to open-world scenarios.

1 papers0 benchmarksTexts, Videos

MuSoHu (Toward human-like social robot navigation: A large-scale, multi-modal, social human navigation dataset)

A large-scale, egocentric, multimodal, and context-aware dataset of human demonstrations of social navigation.

1 papers0 benchmarks3D, Actions, LiDAR, Point cloud, RGB-D, Stereo, Videos

INSPIRE-AVR (LUNet subset)

This dataset contains 65 DFIs acquired from patients with POAG at the University of Iowa Hospitals and Clinics. DFIs were acquired using a 30° Zeiss fundus camera (Niemeijer et al 2011). The images were centered on the optic disc. The original DFIs resolution was 2392 × 2048. In order to benchmark LUNet on this dataset, the black border of the DFIs were padded to a squared resolution of 2048 × 2048 pixels and then resized to a 1444 × 1444 pixels resolution. From the resulting DFIs, 15 optic disc-centered DFIs were randomly selected to form the second external test set. No other additional metadata were provided in the open source dataset.

1 papers4 benchmarksImages

FQ-160 (Forbidden Question Dataset (160))

The forbidden question dataset they build (based on two previous works) contains 160 questions from 160 violated categories. In addition, they also provide the corresponding targets – which is useful for some jailbreak methods, such as GCG.

1 papers0 benchmarks

HALvest

HALvest is a textual dataset comprising 17 billion tokens in 56 languages and 13 domains.

1 papers0 benchmarksTexts

HALvest-Geometric

HALvest-Geometric is a subset of HALvest: an academic citation network with 238,397 disambiguated authors and 18,662,037 scholarly papers.

1 papers0 benchmarksGraphs, Texts

Bluesky Social Dataset

Bluesky Social Dataset Pollution of online social spaces caused by rampaging d/misinformation is a growing societal concern. However, recent decisions to reduce access to social media APIs are causing a shortage of publicly available, recent, social media data, thus hindering the advancement of computational social science as a whole. To address this pressing issue, we present a large, high-coverage dataset of social interactions and user-generated content from Bluesky Social.

1 papers0 benchmarks

HePIC 🏛️

Heritage Pointcloud Instance Collection dataset, acquired from two large buildings and annotated at a point-wise semantic level based on existent BIM models. Devid Campagnolo, Elena Camuffo, Umberto Michieli, Paolo Borin, Simone Milani and Andrea Giordano, "Fully Automated Scan-to-BIM via Point Cloud Instance Segmentation", In Proceedings of the International Conference on Image Processing (ICIP) 2023.

1 papers2 benchmarks3D, Point cloud

BTAT (Blockchain Transaction-based Attacks dataset)

The Synthesis Blockchain Intrusion Detection System dataset

1 papers0 benchmarks

FinSen

Enhancing Financial Market Predictions: Causality-Driven Feature Selection This paper introduces FinSen dataset that revolutionizes financial market analysis by integrating economic and financial news articles from 197 countries with stock market data. The dataset’s extensive coverage spans 15 years from 2007 to 2023 with temporal information, offering a rich, global perspective 160,000 records on financial market news. Our study leverages causally validated sentiment scores and LSTM models to enhance market forecast accuracy and reliability.

1 papers2 benchmarks

Forbidden_Questions_160

Here is the forbidden question dataset (based on two previous works), it contains 160 questions from 160 violated categories. In addition, authors also provide the corresponding target -- which is useful for some jailbreak methods, such as GCG.

1 papers0 benchmarks

Forbidden_Questions_CJA (Forbidden Question Dataset (160))

Here is the forbidden question dataset (based on two previous works), it contains 160 questions from 160 violated categories. In addition, authors also provide the corresponding target -- which is useful for some jailbreak methods, such as GCG.

1 papers0 benchmarks

Expository Prose (Expository-Prose-V1)

Expository-Prose-V1 is a collection of specially-curated corpora gathered from diverse sources, ranging from research papers (arXiv) to European Parliament proceedings (EuroParl). It has been specially filtered and curated for the quality of text, depth of reasoning and breadth of knowledge to faciliate an effective pre-train. It was used to pre-train 1.5-Pints, a small but powerful Large Language Model developed by the Pints Research Team.

1 papers0 benchmarksTexts

Noise of Web (NoW)

Noise of Web (NoW) is a challenging noisy correspondence learning (NCL) benchmark for robust image-text matching/retrieval models. It contains 100K image-text pairs consisting of website pages and multilingual website meta-descriptions (98,000 pairs for training, 1,000 for validation, and 1,000 for testing). NoW has two main characteristics: without human annotations and the noisy pairs are naturally captured. The source image data of NoW is obtained by taking screenshots when accessing web pages on mobile user interface (MUI) with 720 $\times$ 1280 resolution, and we parse the meta-description field in the HTML source code as the captions. In NCR (predecessor of NCL), each image in all datasets were preprocessed using Faster-RCNN detector provided by Bottom-up Attention Model to generate 36 region proposals, and each proposal was encoded as a 2048-dimensional feature. Thus, following NCR, we release our the features instead of raw images for fair comparison. However, we can not just

1 papers0 benchmarksImages, Texts

TML1M (Table-MovieLens1M)

Table-MovieLens1M (TML1M) is a relational table dataset derived from the classical MovieLens1M dataset. It consists of three tables: users, movies, and ratings. Notably, the movie table has been enriched with more comprehensive features. Additionally, the dataset defines a standard classification task focused on predicting user age ranges.

1 papers1 benchmarksTabular

TLF2K (Table-LastFm2K)

Table-LastFm2K (TLF2K) is a relational table dataset derived from the classical LastFM2K dataset. It contains three tables: artists, user_artists, and user_friends. Notably, the artists table has been enhanced with more detailed features, and the tags for each artist have been streamlined. The dataset also provides a standard classification task for music genre classification of artists.

1 papers1 benchmarksTabular

TACM12K (Table-ACM12K)

Table-ACM12K (TACM12K) is a relational table dataset derived from the ACM heterogeneous graph dataset. It includes four tables: papers, authors, citations, and writings. The paper table features attributes such as year, title, and abstract, while the author table includes name and affiliation details. Additionally, some feature completion has been performed for the papers. The dataset also defines a standard classification task for predicting the conference to which a paper belongs.

1 papers1 benchmarksTabular

Hamlyn Dataset

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

Flight Delay and Cancellation Dataset (2019-2023)

The dataset used in this study is obtained from the U.S. Department of Transportation, Bureau of Transportation Statistics from January 2019– August 2023 . It contains 32 attributes related to planned f light date-time, airline, planned origin and destination, cancellation and diversion status, overall delay, and delay due to individual components (carrier, weather, NAS, security, late aircraft), among others.

1 papers0 benchmarks

MedTrinity-25M

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksBiomedical, Images, MRI, Medical, Texts
PreviousPage 512 of 1000Next