TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Biodenoising_validation

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksAudio

Biodenoising datasets

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

CEMS-W (CEMS Wildires)

The dataset includes annotations for burned area delineation and land cover segmentation, with a focus on European soil. The dataset is curated from various sources, including the Copernicus European Monitoring System (EMS) and Sentinel-2 feeds.

1 papers2 benchmarksImages

IVM-Mix-1M

IVM-Mix-1M provide over 1M image-instruction pairs with corresponding instruction-relevant mask labels. Our IVM-Mix-1M dataset consists of three part: HumanLabelData, RobotMachineData and VQAMachineData. For the HumanLabelData and RobotMachineData, we provide well-orgnized images, mask label and language instructions. For the VQAMachineData, we only provide mask label and language instructions, please refer to https://huggingface.co/datasets/2toINF/IVM-Mix-1M and download the images from constituting datasets.

1 papers0 benchmarksImages, Texts

Business License

Business license datasets and source code for named entity recognition.

1 papers0 benchmarks

ListUltraFeedback

A listwise multi-response dataset for human preferences alignment. The dataset is derived from UltraFeedback and SimPO.

1 papers0 benchmarksTexts

SPADA Dataset

Dataset for Land Cover segmentation from sparse labels, using Sentinel-2 as source imagery.

1 papers0 benchmarksImages

Herbarium-19

The FGVC 2019 Herbarium Challenge is to identify melastome species from herbarium specimens provided by the New York Botanical Garden (NYBG). We have provided a curated dataset of over 46,000 herbarium specimens for over 680 species of the flowering plant family Melastomataceae.

1 papers0 benchmarks

MIE Articles Dataset (1996-2024)

This dataset contains 4606 articles from 1996 to 2024 that were presented in MIE (Medical Informatics Europe Conference) conferences. This data was extracted from PubMed and topic extraction and affiliation parsing were done on it.

1 papers0 benchmarksTexts

arXiv Categories (arXiv Categories Multi-label Text Classification Dataset)

This is a dataset of scientific documents derived from arXiv. It comprises 203,961 titles and abstracts categorized into 130 different classes from the arXiv category taxonomy. Each document (title+abstract) is categorized into one or more distinct classes. It is split into train (163,168), validation (20,396), and test (20,397) sets.

1 papers0 benchmarksTexts

Real Estate (CC)

This dataset represents residential real estate listings with the following features:

1 papers0 benchmarks

MolQA

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

TruthGen

TruthGen is a dataset of generated true and false statements, intended for research on truthfulness in reward models and language models, specifically in contexts where political bias is undesirable. This dataset contains 1,987 statement pairs (3,974 statements in total), with each pair containing one objectively true statement and one false statement. It spans a variety of everyday and scientific facts, excluding politically charged topics to the greatest extent possible. The dataset is particularly useful for evaluating reward models trained for alignment with truth, as well as for research on mitigating political bias while improving model accuracy on truth-related tasks.

1 papers0 benchmarksTexts

FewEvent

S. Deng, N. Zhang, J. Kang, Y. Zhang, W. Zhang, and H. Chen, “Meta-learning with dynamic-memory-based prototypical network for few-shot event detection”, in Proceedings of the 13th International Conference on Web Search and Data Mining. WSDM ’20, 2020, p. 151–159.

1 papers0 benchmarks

WSJ POS

M. P. Marcus, B. Santorini, and M. A. Marcinkiewicz, “Building a large annotated corpus of English: The Penn Treebank”, Computational Linguistics, vol. 19, no. 2, pp. 313–330, 1993. [Online]. Available: https://aclanthology.org/J93-2004

1 papers1 benchmarks

sPBC (Super Parallel Bible Corpus)

This is a super-parallel Bible corpus containing 1401 language labels (language_script pairs), meaning that for each verse, we have the translation in other languages. However, this results in a relatively low number of verses, with the current version supporting 103 verses per language.

1 papers0 benchmarksTexts

Dataset for Datacube segmentation via Deep Spectral Clustering"

Synthetic Datasets for ICSC Flagship 2.6.1. "Fast Extended Computer Vision" paper #1 "Datacube segmentation via Deep Spectral Clustering"

1 papers0 benchmarks

vqa-nle-llava

VQA NLE synthetic dataset, made with LLaVA-1.5 using features from GQA dataset. Total number of unique datas: 66684

1 papers0 benchmarksImages, Texts

YesBut

YesBut Dataset (https://yesbut-dataset.github.io) Understanding satire and humor is a challenging task for even current Vision-Language models. In this paper, we propose the challenging tasks of Satirical Image Detection (detecting whether an image is satirical), Understanding (generating the reason behind the image being satirical), and Completion (given one half of the image, selecting the other half from 2 given options, such that the complete image is satirical) and release a high-quality dataset YesBut, consisting of 2547 images, 1084 satirical and 1463 non-satirical, containing different artistic styles, to evaluate those tasks. Each satirical image in the dataset depicts a normal scenario, along with a conflicting scenario which is funny or ironic. Despite the success of current Vision-Language Models on multimodal tasks such as Visual QA and Image Captioning, our benchmarking experiments show that such models perform poorly on the proposed tasks on the YesBut Dataset in Zero-Sh

1 papers0 benchmarksImages, Texts

FQL-Driving

FQL-driving

1 papers2 benchmarks
PreviousPage 522 of 1000Next