TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

SolProp

The SolProp dataset is a valuable resource for predicting the solubility limits of organic solutes in various solvents and at different temperatures. It estimates the solubility of neutral organic molecules in both water and organic solvents across a wide temperature range. The model considers solvation-free energy, solvation enthalpy, Abraham solute parameters, and aqueous solubility at 298K.

0 papers0 benchmarks

TFH_Annotated_Dataset (Thin_Film_head_relevant_Patent_Annotated_Dataset)

Dataset Introduction TFH_Annotated_Dataset is an annotated patent dataset pertaining to thin film head technology in hard-disk. To the best of our knowledge, this is the second labeled patent dataset public available in technology management domain that annotates both entities and the semantic relations between entities, the first one is [1].

0 papers0 benchmarksTexts

SMAC+ Def_infantry_episodic

SMAC+ defensive infantry scenario with sequential episodic buffer

0 papers0 benchmarks

QPT (Quantum Process Tomography)

Quantum process tomography (QPT) is a method for experimentally reconstructing the quantum channel from measurement data. A QPT experiment prepares multiple input states, evolves them by the circuit, then performs multiple measurements in different measurement bases.

0 papers0 benchmarks

Simulacra Aesthetic Captions

Simulacra Aesthetic Captions is a dataset of over 238000 synthetic images generated with AI models such as CompVis latent GLIDE and Stable Diffusion from over forty thousand user submitted prompts. The images are rated on their aesthetic value from 1 to 10 by users to create caption, image, and rating triplets. In addition to this each user agreed to release all of their work with the bot: prompts, outputs, ratings, completely public domain under the CC0 1.0 Universal Public Domain Dedication. The result is a high quality royalty free dataset with over 176000 ratings that can be used for projects such as:

0 papers0 benchmarksImages

AiTLAS: Benchmark Arena

AiTLAS: Benchmark Arena is an open-source benchmark framework for evaluating state-of-the-art deep learning approaches for image classification in Earth Observation (EO).

0 papers0 benchmarksImages

Parkour-dataset (LAAS Parkour dataset)

The LAAS Parkour dataset contains 28 RGB videos capturing human subjects performing four typical parkour techniques: safety-vault, kong vault, pull-up and muscle-up. These are highly dynamic motions with rich contact interactions with the environment. The dataset is provided with the ground truth 3D positions of 16 pre-defined human joints, together with the contact forces at the human subjects' hand and foot joints exerted by the environment.

0 papers0 benchmarks

UMass Citation Field Extraction

The University of Massachusetts Amherst citation field extraction dataset contains labels and segments for extracted citations from articles found on arXiv. Compared to previous standard datasets in citation field extraction, this one had 4 times more data and provided detailed nested labels rather than coarse-grained flat labels, alongside drawing from 4 different academic disciplines versus 1 - namely computer science, mathematics, physics, and quantitative biology.

0 papers0 benchmarksTexts

Real-time Election Results: Portugal 2019 Data Set

Data Set Information:

0 papers0 benchmarks

DNA mutations

In bioinformatics, the issue of mutation discovery and type determination remains a significant concern. The problem is divided by the researchers into binary classification and multi-class problems. When the user wants to know if the DNA sequence has been altered, the issue is a binary classification problem. When it is desirable to identify the principal class of mutation or its sub-classes, the problem becomes more challenging. The primary classes of mutations are deletion, insertion, and replacement mutation, and their sub-classes are (deletion frameshift, deletion in-frame, insertion frameshift, insertion in-frame, silent, missense, nonsense, and read-through). Additionally, answers to sporadic issues like the DNA sequence alignment challenge are necessary for mutation detection techniques. Due to the scarcity of labeled databases, this data set was created by addressing an unlabeled database and creating random mutations of all kinds for the purpose of benefiting from them by res

0 papers0 benchmarks

KID-F (K-pop Idol Dataset - Female)

Description K-pop Idol Dataset - Female (KID-F) is the first dataset of K-pop idol high quality face images. It consists of about 6,000 high quality face images at 512x512 resolution and identity labels for each image.

0 papers0 benchmarksImages

BACC-18 (Bengali Authorship Classification Corpus-18)

The developed BACC-18 contains the text of 18 famous authors of Bengali literature. To build this corpus, we crawled texts from four online sources namely NLTR society for natural language technology research [36], Ebanglalibrary [37], Git repository [38] and Blogs [39]–[40][41]. The maximum number of texts (13,308) are collected from NLTR source whereas minimum number of texts (240) are crawled from Blogs. A self-built automatic web crawler3 is used to scrapping the data from four sources. Due to HTML page structure variation of sources, we used various web crawler instead of a typical crawler. In particular, the proposed research has developed 31 Python crawler which can automatically crawl textual data based on the robots.txt policy. The robots.txt policy ensures the search engine whether a crawler can or cannot crawl the particular text contents from a source.4 Initially, we manually selected the famous and authentic web portal’s hyperlink to collect the author’s texts. Web crawler

0 papers0 benchmarks

Wendi (孙文迪)

Null

0 papers0 benchmarks

RSCD: Large-scale Road Surface Classification Dataset for Autonomous Vehicles

The preview of the road surface states is essential for improving the safety and the ride comfort of autonomous vehicles. This dataset consists of 1 million (240 x 360 pixels) road surface images captured under a wide range of road and weather conditions in China. The original pictures are acquired with a vehicle-mounted camera and then the patches containing only the road surface area are cropped. The images are classified into 27 categories, containing both the friction level, material, and unevenness properties. The dataset is divided into train-set(~960k samples), validation-set(~20k samples), test-set(~50k samples) . This large-scale dataset is useful for developing vision-based road sensing modules to improve the performance of the driving assistance systems.

0 papers0 benchmarks

Semeion (Semeion Handwritten Digit Data Set)

1593 handwritten digits from around 80 persons were scanned, stretched in a rectangular box 16x16 in a gray scale of 256 values.

0 papers0 benchmarksImages

Infinity Spills Basic Dataset

Infinity AI's Spills Basic Dataset is a synthetic, open-source dataset for safety applications. It features 150 videos of photorealistic liquid spills across 15 common settings. Spills take on in-context reflections, caustics, and depth based on the surrounding environment, lighting, and floor. Each video contains a spill of unique properties (size, color, profile, and more) and is accompanied by pixel-perfect labels and annotations. This dataset can be used to develop computer vision algorithms to detect the location and type of spill from the perspective of a fixed camera.

0 papers0 benchmarksImages, RGB Video, Videos

Coronavirus (COVID-19) Tweets Dataset

This dataset includes CSV files that contain IDs and sentiment scores of the tweets related to the COVID-19 pandemic. The real-time Twitter feed is monitored for coronavirus-related tweets using 90+ different keywords and hashtags that are commonly used while referencing the pandemic. The oldest tweets in this dataset date back to October 01, 2019. This dataset has been wholly re-designed on March 20, 2020, to comply with the content redistribution policy set by Twitter. Twitter's policy restricts the sharing of Twitter data other than IDs; therefore, only the tweet IDs are released through this dataset. You need to hydrate the tweet IDs in order to get complete data.

0 papers0 benchmarksTexts

Sparse LiDAR KITTI dataset

Sparse LiDAR extracted from velodyne 64 beams in KITTI dataset. It contains severals LiDAR: LiDAR 2 beams, LiDAR 4 beams, LiDAR 8 beams, LiDAR 16 beams, LiDAR 32 beams

0 papers0 benchmarksLiDAR, Point cloud

Car datasets in multiple scenes

This dataset is a collection of 4,000 images of cars in multiple scenes that are ready to use for optimizing the accuracy of computer vision models. All of the contents is sourced from PIXTA's stock library of 100M+ Asian-featured images and videos. PIXTA is the largest platform of visual materials in the Asia Pacific region offering fully-managed services, high quality contents and data, and powerful tools for businesses & organizations to enable their creative and machine learning projects. For more details, please refer to the link: https://www.pixta.ai/ Or send your inquiries to contact@pixta.ai

0 papers0 benchmarks

Human faces with mixed-race & various emotions

This dataset consists of 600+ items of faces with different emotions and mixed races that are ready to use for optimizing the accuracy of computer vision models. Human age range from 20 to 60 years old, balance in gender, no occlusion, with head direction (<45 degree up-down and left-right). All of the contents is sourced from PIXTA's stock library of 100M+ Asian-featured images and videos. PIXTA is the largest platform of visual materials in the Asia Pacific region offering fully-managed services, high quality contents and data, and powerful tools for businesses & organizations to enable their creative and machine learning projects. For more details, please refer to the link: https://www.pixta.ai/ Or send your inquiries to contact@pixta.ai

0 papers0 benchmarks
PreviousPage 648 of 1000Next