TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Automated Evolution of Feature Logging Statement Levels Using Git Histories and Degree of Interest

Logging—used for system events and security breaches to more informational yet essential aspects of software features—is pervasive. Given the high transactionality of today's software, logging effectiveness can be reduced by information overload. Log levels help alleviate this problem by correlating a priority to logs that can be later filtered. As software evolves, however, levels of logs documenting surrounding feature implementations may also require modification as features once deemed important may have decreased in urgency and vice-versa. We present an automated approach that assists developers in evolving levels of such (feature) logs. The approach, based on mining Git histories and manipulating a degree of interest (DOI) model, transforms source code to revitalize feature log levels based on the "interestingness" of the surrounding code. Built upon JGit and Mylyn, the approach is implemented as an Eclipse IDE plug-in and evaluated on 18 Java projects with ~3 million lines of co

1 papers0 benchmarks

Medical Wiki Paralell Corpus for Medical Text Simplification

A medical Wiki paralell corpus for medical text simplification.

1 papers0 benchmarksTexts

Infologic sql queries

Sql queries

1 papers0 benchmarks

5DOF GB Interpolation (Five Degree-of-Freedom Grain Boundary Interpolation)

These are larger MATLAB .mat files required for reproducing plots from the sgbaird-5DOF/interp repository for grain boundary property interpolation. gitID-0055bee_uuID-475a2dfd_paper-data6.mat contains multiple trials of five degree-of-freedom interpolation model runs for various interpolation schemes. gpr46883_gitID-b473165_puuID-50ffdcf6_kim-rng11.mat contains a Gaussian Process Regression model trained on 46883 Fe simulation GBs. See Five degree-of-freedom property interpolation of arbitrary grain boundaries via Voronoi fundamental zone framework DOI: 10.1016/j.commatsci.2021.110756 for the peer-reviewed, published version of the paper.

1 papers0 benchmarksTabular

Self-stimulatory Behavior Dataset

Autism Spectrum Disorders (ASD), often referred to as autism, are neurological disorders characterised by deficits in cognitive skills, social and communicative behaviours. A common way of diagnosing ASD is by studying behavioural cues expressed by the children.

1 papers0 benchmarks

Waste Classification data

PROBLEM Waste management is a big problem in our country. Most of the wastes end up in landfills. This leads to many issues like: Increase in landfills, Eutrophication, Consumption of toxic waste by animals, Leachate, Increase in toxins, Land, water and air pollution.

1 papers0 benchmarks

DRKG (Drug Repurposing Knowledge graph)

Drug Repurposing Knowledge Graph (DRKG) is a comprehensive biological knowledge graph relating genes, compounds, diseases, biological processes, side effects and symptoms. DRKG includes information from six existing databases including DrugBank, Hetionet, GNBR, String, IntAct and DGIdb, and data collected from recent publications particularly related to Covid19. It includes 97,238 entities belonging to 13 entity-types; and 5,874,261 triplets belonging to 107 edge-types. These 107 edge-types show a type of interaction between one of the 17 entity-type pairs (multiple types of interactions are possible between the same entity-pair), as depicted in the figure below. It also includes a bunch of notebooks about how to explore and analysis the DRKG using statistical methodologies or using machine learning methodologies such as knowledge graph embedding.

1 papers0 benchmarks

CorruptionDataSet

This original data set includes the following four sheets: Sheet 1: Raw Data (the original data set) Sheet 2: Variables (A list with the variables included in the study) Sheet 3: Countries Scientific Relative Production Sheet 4: Correlations

1 papers0 benchmarks

Helix

See https://zenodo.org/record/5500215#.YUCgD51Kg2w

1 papers0 benchmarks

Dataset of 3D Garments with Sewing Patterns

The Dataset contains more than 23500 3D garment models with their corresponding sewing patterns, each representing a unique garment design sampled from one of the 19 different categories. The dataset is suitable for training Deep Learning models to solve a variety of clothing-related tasks.

1 papers0 benchmarks

YorkTag

YorkTag provides pairs of sharp/blurred images containing fiducial markers and is proposed to train and qualitatively and quantitatively evaluate our model.

1 papers0 benchmarksImages

E-Manual Corpus

E-Manual Corpus is a corpus of 307,957 E-manuals, used for pre-training models for Question Answering on e-manuals.

1 papers0 benchmarksTexts

DMO

A large scale dataset to pre-train optical flow prediction network. The data are generated from the DAVIS videos using as-rigid-as-possible principle from Deep-matching and MaskRCNN. The dataset has shown better performance compared to the FlyingChairs dataset.

1 papers0 benchmarks

Nelson-Plosser (Nelson-Plosser US Macroeconomic Time Series)

US Macroeconomic dataset containing 14 time series of monthly observations. They have various lengths but all end in 1988. The variables: consumer price index, industrial production, nominal GNP, velocity, employment, interest rate, nominal wages, GNP deflator, money stock, real GNP, stock prices (S&P500), GNP per capita, real wages, unemployment.

1 papers0 benchmarksTables, Time series

BLANCA

BLANCA (Benchmarks for LANguage models on Coding Artifacts) is a collection of benchmarks that assess code understanding based on tasks such as predicting the best answer to a question in a forum post, finding related forum posts, or predicting classes related in a hierarchy from class documentation.

1 papers0 benchmarksTexts

ELITR ECA

The ELITR ECA corpus is a multilingual corpus derived from publications of the European Court of Auditors. We use automatic translation together with Bleualign to identify parallel sentence pairs in all 506 translation directions. The result is a corpus comprising 264k document pairs and 41.9M sentence pairs.

1 papers0 benchmarksTexts

EDGAR10-Q Dataset

This dataset is built from 10-Q documents (Quarterly Reports) of publicly listed companies on the SEC.

1 papers0 benchmarks

wikiHow-image

The dataset consists of 53,189 wikiHow articles across various categories of everyday tasks, 155,265 methods, and 772,294 steps with corresponding images.

1 papers1 benchmarksImages, Texts

Depth VIDIT (Virtual Image Dataset for Illumination Transfer)

VIDIT is a reference evaluation benchmark and to push forward the development of illumination manipulation methods. Virtual datasets are not only an important step towards achieving real-image performance but have also proven capable of improving training even when real datasets are possible to acquire and available. VIDIT contains 300 virtual scenes used for training, where every scene is captured 40 times in total: from 8 equally-spaced azimuthal angles, each lit with 5 different illuminants.

1 papers0 benchmarks

ADEFAN

This data set contains 50 low resolution (640 x 360) short videos containing a variety real life activities.

1 papers0 benchmarks
PreviousPage 406 of 1000Next