TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

CCIC (Concrete Crack Images for Classification)

The dataset contains concrete images having cracks. The data is collected from various METU Campus Buildings. The dataset is divided into two as negative and positive crack images for image classification. Each class has 20000images with a total of 40000 images with 227 x 227 pixels with RGB channels. The dataset is generated from 458 high-resolution images (4032x3024 pixel) with the method proposed by Zhang et al (2016). High-resolution images have variance in terms of surface finish and illumination conditions. No data augmentation in terms of random rotation or flipping is applied.

0 papers0 benchmarks

RAISE-LPBF

Laser powder bed fusion (LBPF) is the additive manufacturing (3D printing) process for metals. RAISE-LPBF is a large dataset on the effect of laser power and laser dot speed in 316L stainless steel bulk material. Both process parameters are independently sampled for each scan line from a continuous distribution, so interactions of different parameter choices can be investigated. Process monitoring comprises on-axis high-speed (20k FPS) video. The data can be used to derive statistical properties of LPBF, as well as to build anomaly detectors.

0 papers0 benchmarksImages, Physics, Videos

AudioSet CC (AudioSet Creative Commons)

The subset of audio samples from the AudioSet ontology which are licensed with Creative Commons. This set contains approximately 10,000 samples of 10s long clips, and is freely modifiable and distributable. Each clip has with it, its full label set and unique ID.

0 papers0 benchmarksAudio

Replication Data for: AI Ethics on Blockchain: Topic Analysis on Twitter Data for Blockchain Security

Blockchain has empowered computer systems to be more secure using a distributed network. However, the current blockchain design suffers from fairness issues in transaction ordering. Miners are able to reorder transactions to generate profits, the so-called miner extractable value (MEV). Existing research recognizes MEV as a severe security issue and proposes potential solutions, including prominent Flashbots. However, previous studies have mostly analyzed blockchain data, which might not capture the impacts of MEV in a much broader AI society. Thus, in this research, we applied natural language processing (NLP) methods to comprehensively analyze topics in tweets on MEV. We collected more than 20000 tweets with #MEV and #Flashbots hashtags and analyzed their topics. Our results show that the tweets discussed profound topics of ethical concern, including security, equity, emotional sentiments, and the desire for solutions to MEV. We also identify the co-movements of MEV activities on blo

0 papers0 benchmarks

Replication Data for: On the Mechanics of NFT Valuation: AI Ethics and Social Media

As CryptoPunks pioneers the innovation of non-fungible tokens (NFTs) in AI and art, the valuation mechanics of NFTs has become a trending topic. Earlier research identifies the impact of ethics and society on the price prediction of CryptoPunks. Since the booming year of the NFT market in 2021, the discussion of CryptoPunks has propagated on social media. Still, existing literature hasn't considered the social sentiment factors after the historical turning point on NFT valuation. In this paper, we study how sentiments in social media, together with gender and skin tone, contribute to NFT valuations by an empirical analysis of social media, blockchain, and crypto exchange data. We evidence social sentiments as a significant contributor to the price prediction of CryptoPunks. Furthermore, we document structure changes in the valuation mechanics before and after 2021. Although people's attitudes towards Cryptopunks are primarily positive, our findings reflect imbalances in transaction act

0 papers0 benchmarks

The Rambles

Collection of stream of consciousness.

0 papers0 benchmarks

Inria building dataset

Inria building dataset contains 360 images (5120×5120) collected from 5 cities (Austin, Chicago, Kitsap, Tyrol, and Vienna)

0 papers0 benchmarks

MoralChoice Survey

MoralChoice is a survey dataset to evaluate the moral beliefs encoded in LLMs. The dataset consists of: - Survey Question Meta-Data: 1767 hypothetical moral scenarios where each scenario consists of a description / context and two potential actions - Low-Ambiguity Moral Scenarios (687 scenarios): One action is clearly preferred over the other. - High-Ambiguity Moral Scenarios (680 scenarios): Neither action is clearly preferred - Survey Question Templates: 3 hand-curated question templates - Survey Responses: Outputs from 28 open- and closed-sourced LLMs

0 papers0 benchmarksTexts

CIDII Dataset (Correct Information and Disinformation about Islamic Issues)

The CIDII dataset is a binary classification, consisting of two classes of correct information and disinformation related to Islamic issues. The CIDII dataset belongs to our research (DISINFORMATION DETECTION ABOUT ISLAMIC ISSUES ON SOCIAL MEDIA USING DEEP LEARNING TECHNIQUES) published in MJCS journal in the link below: https://ejournal.um.edu.my/index.php/MJCS/article/view/41935

0 papers0 benchmarksTexts

Volumetric CMR Cartesian Datasets (Free-running self-gated 3D cine, 4D Flow and stress 4D Flow Undersampled Datasets)

Datasets at https://zenodo.org/record/8105485 for Motion Robust CMR Reconstruction Code in https://github.com/syedmurtazaarshad/motion-robust-CMR

0 papers0 benchmarksBiomedical, Images, MRI

ALTA 2023 Shared Task (Discriminate between human-authored and synthetic text generated by Large Language Models (LLMs))

This dataset is described in the ALTA 2023 Shared Task and associated CodaLab competition.

0 papers0 benchmarksTexts

ALTA 2022 Shared Task (PIBOSO Sentence classification)

This dataset is described in the ALTA 2022 Shared Task and associated CodaLab competition.

0 papers0 benchmarksTexts

Raw_-Subjective-Scores-120-videos

A review on raw subjective scores and data manipulation for before and after refining Mean opinion Scores

0 papers0 benchmarksTexts

VIDIMU: Multimodal video and IMU kinematic dataset on daily life activities using affordable devices (https://zenodo.org/record/8210563)

Human activity recognition and clinical biomechanics are challenging problems in physical telerehabilitation medicine. However, most publicly available datasets on human body movements cannot be used to study both problems in an out-of-the-lab movement acquisition setting. The objective of the VIDIMU dataset is to pave the way towards affordable patient tracking solutions for remote daily life activities recognition and kinematic analysis.

0 papers0 benchmarks3D, Biomedical, RGB Video, Time series, Videos

Accompanying dataset for: Predicting Species Emergence in Simulated Complex Pre-Biotic Networks

This is the accompanying data and code for the publication [Markovitch & Krasnogor: Predicting Species Emergence in Simulated Complex Pre-Biotic Networks] containing the full set of 10,000 lognormal networks studied, their network communities and the compotype species observed during simulations with the GARD model. Details are given in the aforementioned paper.

0 papers0 benchmarks

WRV (Wire-removal Dataset)

G2LP Wire-removal Dataset in G2LP-Net: Global to Local Progressive Video Inpainting Network

0 papers0 benchmarks

Bomstic

Plant growth

0 papers0 benchmarks

Text_VPH

Este conjunto de datos consiste en comentarios de publicaciones del MINSA (Perú) en Facebook sobre la vacuna contra el VPH entre los años 2019 y 2020. Se leyó cuidadosamente cada uno de los comentarios, luego se procedió a clasificarlos de manera manual. Para esta clasificación se interpretó los mensajes de las personas, por lo que se analizó los hilos (comentarios y respuestas) por separado y se procedió a etiquetarlos por temas "Topic" . Un profesional de salud realizó una segunda clasificación y las discrepancias se resolvieron con un tercer profesional. Luego, se seleccionaron subcategorías que hacían referencia directa a las vacunas contra el VPH. La clasificación se realizó utilizando las siguientes categorías "topic_c" :

0 papers0 benchmarks

FinArg (Financial Argument Mining)

With the goal of reasoning on the financial textual data, we present a novel dataset for annotating arguments, their components, and relations in the transcripts of earnings conference calls (ECCs).

0 papers0 benchmarks

Chaotic Trajectories

Dataset of chaotic Chua, Lorenz, Lorenz96, Mackey-Glass with tau=17, Mackey-Glass with tau=30, Rossler, Sprott systems.

0 papers0 benchmarks
PreviousPage 653 of 1000Next