TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Pedestrian yielding

The data collection took place at 18 intersections in Minnesota. Data collection was conducted at 18 sites across Minnesota to capture various road types, and intersection features during the summer months. The TIMs were deployed to each location for between one and two weeks, and video was recorded during daylight hours only. The Minnesota Traffic Observatory (MTO) used the Traffic Information Monitors (TIMs) to collect video data at each site. Raw video data will be available upon request. The collected video data is reduced to numeric data using both video processing techniques and manual data extraction.

0 papers0 benchmarks

Fields2Benchmark dataset

The Fields2Benhmark dataset is a collection of 350 agricultural fields in vector format manually selected to test agricultural coverage path planning algorithms.

0 papers0 benchmarksEnvironment, Graphs, Images

HkgEA

123

0 papers0 benchmarks

test10

test10

0 papers0 benchmarks

烟草茎叶角 (茎叶角检测数据集)

烟草茎叶角

0 papers0 benchmarks

English-Pashto Language Dataset (EPLD)

The English-Pashto Language Dataset (EPLD) is a comprehensive resource aimed to provide linguistic insights into the Pashto language. It contains the knowledge and study of Pashto language with the basics of communication like counting, alphabets, pronoun, basic sentences used in everyday life. Every data is translated from English to Pashto for better human understanding and clarity. The data is carefully proofread and verified by the native speakers and the language experts. Pashto language has multiple variations and accents depending on the geographical factors. This dataset explains and addresses the key differences of words and sounds of Pashto, which may sound similar or different from English on the basis of gender, tense of the statement, relationship of the speaker etc. This dataset is designed to support language learning, natural language processing (NLP) research and computational linguistic studies focusing on Pashto language.

0 papers0 benchmarksTabular, Texts

french-legal-cases

annotations_creators: - no-annotation language: - fr language_creators: - found license: - cc-by-4.0 multilinguality: - monolingual pretty_name: French Legal Cases Dataset size_categories: - n>1M source_datasets: - la-mousse/INCA-17-01-2025 - la-mousse/JADE-17-01-2025 - la-mousse/CASS-17-01-2025 - la-mousse/CAPP-17-01-2025 task_categories: - text-generation - text-classification task_ids: - language-modeling - entity-linking-classification - fact-checking - intent-classification - multi-label-classification - multi-input-text-classification - natural-language-inference - semantic-similarity-classification - sentiment-classification - topic-classification - sentiment-analysis - named-entity-recognition - parsing - extractive-qa - open-domain-qa - closed-domain-qa - news-articles-summarization - news-articles-headline-generation - dialogue-modeling - dialogue-generation - abstractive-qa - closed-domain-qa - keyword-spotting - semantic-segmentation - tabular-multi-class-classification -

0 papers0 benchmarks

ML4TSP-Uniform

This is the standard dataset for solving the TSP problem with uniformly distributed points using machine learning, covering six scales: 50, 100, 200, 500, 1K, and 10K.

0 papers0 benchmarks

SUIM-E (SGUIE-Net: Semantic attention guided underwater image enhancement with multi-scale perception)

Underwater Image Enhancement Dataset

0 papers0 benchmarks

Concrete Crack and Spall Dataset for Segmentation and Classification

This dataset, which can be used for vision-based deep learning methods, can been implemented to detect and analyze damages in concrete structures. In order to increase the generalizability of the network results, a number of images from three datasets in references [1-3] were combined as follows which can be downloaded from https://www.kaggle.com/datasets/stmlen/cconcrack. (If you use this dataset, please cite this research paper: https://arxiv.org/abs/2501.11836) • One hundred Turkish images were randomly selected from the Özgenel segmentation dataset [1] with a resolution of 30244032. • One hundred Turkish images were randomly selected from the reference dataset [3] with different resolutions such as 296306 and 334306. • Two hundred crack and drop images were randomly selected from the reference dataset [2] with different resolutions such as 768768 and 960*1280. In general, 400 concrete cracks and spall images (2 categories) with different resolutions were randomly selected from the

0 papers0 benchmarks

Cadenza Woodwind

This publicly available data is synthesised audio for woodwind quartets including renderings of each instrument in isolation. The data was created to be used as training data within Cadenza's second open machine learning challenge (CAD2) for the task on rebalancing classical music ensembles. The dataset is also intended for developing other music information retrieval (MIR) algorithms using machine learning. It was created because of the lack of large-scale datasets of classical woodwind music with separate audio for each instrument and permissive license for reuse. Music scores were selected from the OpenScore String Quartet corpus. These were rendered for two woodwind ensembles of (i) flute, oboe, clarinet and bassoon; and (ii) flute, oboe, alto saxophone and bassoon. This was done by a professional music producer using industry-standard software. Virtual instruments were used to create the audio for each instrument using software that interpreted expression markings in the score. Co

0 papers0 benchmarksAudio, Music, Stereo

PatentDesc

Patent Desc

0 papers0 benchmarks

e2nerf

E2NeRF

0 papers0 benchmarks

Gap Pattern Detection (Gap Pattern (Gap Up and Gap Down) Detection in Candlestick Trading Charts for Technical Analysis)

1. Candlestick Charts Candlestick charts are a type of financial chart used to represent the price movement of an asset (e.g., stocks, cryptocurrencies) over time. Each "candlestick" consists of: - Body: Represents the opening and closing prices. - Wicks (or Shadows): Represent the highest and lowest prices during the time period.

0 papers0 benchmarksImages, Videos

Dataset for emotional abuse assessment (Dataset for emotional abuse assessment by Deen Mohd)

The dataset for the Deenz Emotional Abuse Scale (DEAS-18) study comprises detailed participant responses collected to validate the psychometric properties of the scale. It contains 90 observations across multiple variables, including demographic information (e.g., age, gender) and DEAS-18 subscale scores (Belittling, Isolation, and Threatening Behavior). The dataset was compiled from a sample of college students aged 18-30 years, recruited through convenience sampling.

0 papers0 benchmarks

Monopedal Gaits (Periodic Trajectories of a Passive One-Legged Hopper)

The dataset comprises time-series data capturing distinct periodic motions (gaits) of an energetically conservative one-legged hopper. Due to energy conservation, all gaits form a continuous one-dimensional family and undergo bifurcations as the internal energy varies, leading to different motion patterns such as in-place hopping, forward hopping, and backward hopping.

0 papers0 benchmarks

Axial turbine dataset

An axial turbine is a simplest hydrulic machine which is suitable for low-head conditions. It is required to simulate a machine through CFD and post-process the results to compare the performances of different machines. However, performing CFD simulations are computationally time-consuming and expensive. The goal of this repository is to provide the database of axial turbines and their corresponding post-processed results.

0 papers0 benchmarks

depression interview dataset (depression interview dataset with 1.6 million clinical trail data)

contain the clinical trial dataset

0 papers0 benchmarksTabular

Traffic Sign Recognition YOLOv8

Description: The Traffic Sign Recognition Dataset is designed to support the development of deep learning models, particularly for object detection and classification. The dataset includes images of various traffic signs, each annotated with bounding boxes and corresponding class labels. These images have been preprocessed for uniform size and pixel normalization, ensuring optimal training conditions for models like YOLOv8. The dataset captures a wide variety of traffic signs, making it ideal for tasks related to traffic safety and autonomous vehicle systems.

0 papers0 benchmarks

CytoImage Net Dataset

Description: CytoImageNet is an extensive collection of microscopy images, carefully curated to aid in the development of fast and automated methods for analyzing biological data. With over 890,000 grayscale images spanning 894 diverse classes, it addresses the increasing demand for high-throughput image-based biological assays.

0 papers0 benchmarks
PreviousPage 667 of 1000Next