TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

World Across Time

The World Across Time (WAT) dataset used in paper "CLNeRF: Continual Learning Meets NeRF". It contains multiple colmap reconstructed scenes used for continual learning of NeRFs. For each scene, we provide multiple scans captured at different time where the same scene has different appearance and geometry conditions.

1 papers0 benchmarks

TUR2SQL

The field of converting natural language into corresponding SQL queries using deep learning techniques has attracted significant attention in recent years. While existing Text-to-SQL datasets primarily focus on English and other languages such as Chinese, there is a lack of resources for the Turkish language. In this study, we introduce the first publicly available cross-domain Turkish Text-to-SQL dataset, named TUR2SQL. This dataset consists of 10,809 pairs of natural language statements and their corresponding SQL queries. We conducted experiments using SQLNet and ChatGPT on the TUR2SQL dataset. The experimental results show that SQLNet has limited performance and ChatGPT has superior performance on the dataset. We believe that TUR2SQL provides a foundation for further exploration and advancements in Turkish language-based Text-to-SQL research.

1 papers0 benchmarks

DGL Version of OpenCatalyst (OC20) ISRE

We provide DGL compatible graphs in lmdb format for the OpenCatalyst IS2RE task based on the OC20 dataset.

1 papers0 benchmarks

MatSci-NLP Benchmark Dataset

We present MatSci-NLP, a natural language benchmark for evaluating the performance of natural language processing (NLP) models on materials science text. We construct the benchmark from publicly available materials science text data to encompass seven different NLP tasks, including conventional NLP tasks like named entity recognition and relation classification, as well as NLP tasks specific to materials science, such as synthesis action retrieval which relates to creating synthesis procedures for materials.

1 papers0 benchmarks

OCB (Open Circuit Benchmark)

OCB contains two graph datasets, Ckt-Bench-101 and Ckt-Bench-301, for representation learning over analog circuits. Ckt-Bench-101 and Ckt-Bench-301 contain graphs (DAGs) that represent analog circuits and provide their corresponding graph-level properties: DC gain (Gain), bandwidth (BW), phase margin (PM),Figure of Merit (FoM), which characterize the circuit performance.

1 papers0 benchmarks

The HYPSO-1 Sea-Land-Cloud-Labeled Dataset

Hyperspectral Imaging, employed in satellites for space remote sensing, like HYPSO-1, faces constraints due to few labeled data sets, affecting the training of AI models demanding these ground-truth annotations. In this work, we introduce The HYPSO-1 Sea-Land-Cloud-Labeled Dataset, an open dataset with 200 diverse hyperspectral images from the HYPSO-1 mission, available in both raw and calibrated forms for scientific research in Earth observation. Moreover, 38 of these images from different countries include ground-truth labels at pixel-level totaling about 25 million spectral signatures labeled for sea/land/cloud categories. To demonstrate the potential of the dataset and its labeled subset, we have additionally optimized a deep learning model (1D Fully Convolutional Network), achieving superior performance to the current state of the art. Our dataset supports applications like super-resolution, anomaly detection, image fusion, classification, and unmixing. The complete dataset, groun

1 papers0 benchmarks

Gait3D-Parsing

Gait3D-Parsing is a dataset for gait recognition in the wild. It is an extension of the large-scale and challenging Gait-3D dataset which is collected from an in-the-wild environment. The train set has 3,000 IDs, and the test set has 1,000 IDs. Meanwhile, 1,000 sequences in the test set are taken as the query set, and the rest of the test set is taken as the gallery set.

1 papers0 benchmarks3D

Iridium Message Headers (25MS/s) (Dataset for "Watch This Space: Securing Satellite Communication through Resilient Transmitter Fingerprinting")

Labelled dataset of Iridium “ring alert” downlink messages, including message headers captured at 25MS/s. Message metadata includes satellite and transmitter identifier, satellite position, timestamp, and estimated noise level. The dataset contains 1706556 messages.

1 papers0 benchmarks

SatIQ Model Weights (Model Weights for "Watch This Space: Securing Satellite Communication through Resilient Transmitter Fingerprinting")

Model weights for use with the SatIQ fingerprinting models used in the paper “Watch This Space: Securing Satellite Communication through Resilient Transmitter Fingerprinting”. The models are used to authenticate Iridium satellites from high sample rate message headers.

1 papers0 benchmarks

CongNaMul

CongNaMul Dataset

1 papers0 benchmarks

TILT corpus (GDPR machine-readable transparency information powered by the Transparency Information Language and Toolkit)

A corpus of GDPR machine-readable transparency information powered by the Transparency Information Language and Toolkit (TILT). These statements were extracted from real-world services for academic research purposes. They contain information about the collection, processing, and use of personal data in accordance with the legal requirements of the GDPR. The corpus makes it possible to process the information for various applications, such as automated checks or analyses, and to illustrate the practical applicability.

1 papers0 benchmarks

TSTTC

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

ISM Images

Click to add a brief description of the dataset (Markdown and LaTeX enabled). https://github.com/ajaygunalan/BrightEyes-ISM/tree/main/compressive_ism/data Provide:

1 papers0 benchmarks

AI-ready multiplex IHC-IF dataset (AI-ready restained and co-registered multiplex dataset for head-and-neck squamous cell carcinoma)

We introduce a new AI-ready computational pathology dataset containing restained and co-registered digitized images from eight head-and-neck squamous cell carcinoma patients. Specifically, the same tumor sections were stained with the expensive multiplex immunofluorescence (mIF) assay first and then restained with cheaper multiplex immunohistochemistry (mIHC). This is a first public dataset that demonstrates the equivalence of these two staining methods which in turn allows several use cases; due to the equivalence, our cheaper mIHC staining protocol can offset the need for expensive mIF staining/scanning which requires highly skilled lab technicians. As opposed to subjective and error-prone immune cell annotations from individual pathologists (disagreement > 50%) to drive SOTA deep learning approaches, this dataset provides objective immune and tumor cell annotations via mIF/mIHC restaining for more reproducible and accurate characterization of tumor immune microenvironment (e.g. for

1 papers0 benchmarksBiology, Images, Medical

Facial Skeletal angles (Facial Skeletal Angles (Glabella and Maxilla Angle and Length and Width of Piriformis))

Facial Skeletal Angles (Glabella and Maxilla Angle and Length and Width of Piriformis)

1 papers0 benchmarksBiology, Medical

FormAI Dataset

FormAI is a novel AI-generated dataset comprising 112,000 compilable and independent C programs. All the programs in the dataset were generated by GPT-3.5-turbo using dynamic zero-shot prompting technique and comprises programs with varying levels of complexity. Some programs handle complicated tasks such as network management, table games, or encryption, while others deal with simpler tasks like string manipulation. Each program is labelled based on vulnerabilities present in the code using a formal verification method based on the Efficient SMT-based Bounded Model Checker (ESBMC). This strategy conclusively identifies vulnerabilities without reporting false positives (due to the presence of counter examples), or false negatives (up to a certain bound). The labeled samples can be utilized to train Large Language Models (LLMs) since they contain the exact program location of the software vulnerability.

1 papers0 benchmarks

Unity Synthetic Humans

A package for creating Unity Perception compatible synthetic people.

1 papers0 benchmarks

CLPD (China License Plate Dataset)

The CLPD dataset comprises 1200 images that encompass various regions within mainland China. These images were sourced from diverse origins, including the internet, mobile devices, and in-car recording devices. While the majority of the images were recorded during daylight hours, a portion of them were captured at nighttime. The dataset predominantly features passenger cars, with a limited number of images depicting trucks and buses.

1 papers0 benchmarksImages

CD-HARD

CD-HARD comprises 102 images featuring vehicles with oblique license plates sourced from the Cars dataset. Each image within this dataset exclusively depicts a single vehicle and was captured during daylight hours. While the dataset encompasses images from diverse geographic regions, it predominantly consists of images seemingly taken in European locales.

1 papers0 benchmarksImages

CSPRD (Chinese Stock Policy Retrieval Dataset)

The Chinese Stock Policy Retrieval Dataset (CSPRD) contains a Chinese policy corpus of 10,002 articles and 709 prospectus examples from 545 companies listed on China’s Science and Technology Innovation Board (STAR Market). CSPRD is bilingual in Chinese and English (Translated by ChatGPT) and is annotated by experienced experts from Shanghai Stock Exchange.

1 papers0 benchmarksTexts
PreviousPage 471 of 1000Next