TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

WOS Hierarchical Text Classification

The WOS Hierarchical Text Classification are three dataset variants created from Web of Science (WOS) title and abstract data categorised into a hierarchical, multi-label class structure. The aim of the sampling and filtering methodology used was to create well-balanced class distributions (at chosen hierarchical levels). Furthermore, the WOS_JTF variant was also created with the aim to only contain publication data such that their class assignments results is classes instances that semantically more similar.

1 papers0 benchmarksTexts

https://zenodo.org/records/7495109

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

ROVER

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksImages

ArEEG_Words

ArEEG_Words dataset is a novel EEG dataset recorded from 22 participants with mean age of 22 years (5 female, 17 male) using a 14-channel Emotiv Epoc X device. The participants were asked to be free from any effects on their nervous system, such as coffee, alcohol, cigarettes, and so 8 hours before recording. They were asked to stay calm in a clam room during imagining one of the 16 Arabic Words for 10 seconds. The words include 16 commonly used words such as up, down, left, and right. A total of 352 EEG recordings were collected, then each recording was divided into multiple 250ms signals, resulting in a total of 15,360 EEG signals.

1 papers0 benchmarks

ArEEG_Chars

ArEEG_Chars, the first EEG dataset for Arabic characters, consists of high-quality recordings for 31 unique characters from 30 participants (21 males and 9 females) using the Epoc X 14-channel device. Each participant was asked to focus on all Arabic letters for 10-second segments each. Each Folder is named after the letter and it contains 30 CSV file of wave records.

1 papers0 benchmarks

Artificial fluorescent bacteria dataset

These images consist of a series of bacteria of the type Bacillus Subtilis that are suspended and captured by a digital microscope. The fluorescent bacteria dataset can be created as desired, defining the number of bacteria per image and the total number of images. It comes with 3280x2464 resolution images and centroid locations of each bacteria, useful for enumeration or density map estimation.

1 papers0 benchmarksImages

BN-AuthProf (Bangla Author Profiling Dataset)

Although research on author profiling has quite progressed in abundant resources languages, it is still infancy for limited resources languages such as Bengali. This repository contains our Bangla Author Profiling Dataset (BN-AuthProf). The primary objective is to introduce and benchmark the performance of machine learning approaches on Age and Gender Classification tasks from the social media status of people.

1 papers6 benchmarksTexts

Luna-1

This is the supporting dataset for the ECCV 2024 paper "MARs: Multi-view Attention Regularizations for Patch-based Feature Recognition of Space Terrain". It contains 5,067 cropped images of craters on the surface of the Moon, generated in the Blender software using publicly available NASA mission data. Also included are 2,161 replicate orbital navigation frames from real-world Lunar Reconnaissance Orbiter (LRO) spacecraft poses with ground-truth bounding box annotations. Luna-1 supports the evaluation of learning-based vision systems for Terrain Relative Navigation (TRN) spacecraft applications.

1 papers0 benchmarks

PV-VTT

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

Coil100-Augmented

This dataset derives from Coil100. There are more than 1,1M images of 100 objects. Each object was turned on a turnable through 360 degrees to vary object pose with respect to a fixed color camera. Images of the objects were taken at pose intervals of 5 degrees. This corresponds to 72 poses per object. Then planar rotation (9 angles) and 18 scaling factors has been applied. Objects have a wide variety of complex geometric and reflectance characteristics.

1 papers0 benchmarks

MIPD (Manipulation and Intention In a Novel Corpus of Polish Disinformation)

A novel corpus of 15,356 Polish web articles, including articles identified as containing disinformation. Our dataset enables a multifaceted understanding of disinformation. We present a distinctive multilayered methodology for annotating disinformation in texts. What sets our corpus apart is its focus on uncovering hidden intent and manipulation in disinformative content. A team of experts annotated each article with multiple labels indicating both disinformation creators’ intents and the manipulation techniques employed.

1 papers0 benchmarksTexts

EyeDentify++

EyeDentify++, a dataset specifically designed for pupil diameter estimation based on webcam images and enhanced using Super Resolution techniques.

1 papers0 benchmarks

BraTS PEDs 2023 (The Brain Tumor Segmentation (BraTS) Challenge 2023: Focus on Pediatrics (CBTN-CONNECT-DIPGR-ASNR-MICCAI BraTS-PEDs))

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks3D, MRI, Medical

SynoClip

SynoClip Dataset The SynoClip dataset is a comprehensive and standard dataset specifically designed for the video synopsis task. It consists of six videos, ranging from 8 to 45 minutes, captured from outdoor-mounted surveillance cameras. This dataset is annotated with tracking information, making it an ideal resource not only for video synopsis but also for related tasks such as object detection in videos and multi-object tracking.

1 papers0 benchmarksVideos

DEJAN Synthetic Taxonomy Dataset

This dataset contains synthetic text data generated to train models for text generation. The data is created using a large language model (LLM), specifically the GEMMA-2-9B-IT model, and is designed to encompass a diverse range of scenarios and writing styles.

1 papers0 benchmarks

EGO-CH-Gaze (Learning to Detect Attended Objects in Cultural Sites with Gaze Signals and Weak Object Supervision)

To study the problem of weakly supervised attended object detection in cultural sites, we collected and labeled a dataset of egocentric images acquired from subjects visiting a cultural site. The dataset has been designed to offer a snapshot of the subject’s visual experience while visiting a museum and contains labels for several artworks and details attended by the subjects.

1 papers0 benchmarksEnvironment, Images, Videos

SynFinTabs

Table extraction from document images is a challenging AI problem, and labelled data for many content domains is difficult to come by. Existing table extraction datasets often focus on scientific tables due to the vast amount of academic articles that are readily available, along with their source code. However, there are significant layout and typographical differences between tables found across scientific, financial, and other domains. Current datasets often lack the words, and their positions, contained within the tables, instead relying on unreliable OCR to extract these features for training modern machine learning models on natural language processing tasks. Therefore, there is a need for a more general method of obtaining labelled data. We present SynFinTabs, a large-scale, labelled dataset of synthetic financial tables. Our hope is that our method of generating these synthetic tables is transferable to other domains. To demonstrate the effectiveness of our dataset in training

1 papers0 benchmarks

NeLoRa-Bench

Implementation We use the USRP N210 SDR platform for capturing over-the-air LoRa signals, operating on a UBX daughter board at the 470MHz bands and a sampling rate of 1MHz. The captured signal samples are then delivered to a back-end host for pre-processing and demodulation. On the transmitter side, we use SX1278 client radio based commodity LoRa nodes for transmitting LoRa packets.

1 papers0 benchmarks

DrIFT

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksImages

COVID-19 Lung CT Scans (Mehrad Aria)

This is a large public COVID-19 (SARS-CoV-2) lung CT scan dataset, containing total of 8,439 CT scans which consists of 7,495 positive cases (COVID-19 infection) and 944 negative ones (normal and non-COVID-19). Data is available as 512×512px PNG images and have been collected from real patients in radiology centers of teaching hospitals of Tehran, Iran. The aim of this dataset is to encourage the research and development of effective and innovative methods such as deep CNNs which are able to identify if a person is infected by COVID-19 through the analysis of his/her CT scans. As a baseline for this dataset we used a CNN-based approach inspired by transfer learning which we could achieve an accuracy of 99.61% which is very promising.

1 papers0 benchmarks
PreviousPage 532 of 1000Next