TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Bornholmsk

This dataset is parallel text for Bornholmsk and Danish.

1 papers0 benchmarksTexts

bajer_danish_misogyny (Bajer Online Misogyny)

This is a high-quality dataset of annotated posts sampled from social media posts and annotated for misogyny. Danish language.

1 papers2 benchmarksTexts

SHAJ (Spoken Hate in the Albanian Jargon)

This is an abusive/offensive language detection dataset for Albanian. The data is formatted following the OffensEval convention. Data is from Instagram and YouTube comments.

1 papers2 benchmarksTexts

polstance (Political Stance in Danish)

Political stance in Danish. Examples represent statements by politicians and are annotated for, against, or neutral to a given topic/article.

1 papers0 benchmarksTexts

FM WILN (Montmorency Forest WILN dataset)

This dataset was created while conducting the field report related to this paper. It includes 18.6 km of autonomous navigation in a boreal forest. The wintertime meteorological conditions are documented in the paper.

1 papers0 benchmarks

MODA dataset (Massive Online Data Annotation Spindle Dataset)

MODA is a large open-source dataset of high quality, human-scored sleep spindles (5342 spindles, from 180 subjects) that was produced by the Massive Online Data Annotation project. Sleep spindles were detected as a consensus of a number of human-expert scorers. With a median number of 5 experts scoring every EEG segment, MODA offers sleep spindle annotations of a quality unseen in previous datasets.

1 papers3 benchmarksEEG, Medical

SSVC (Synthetic Structured Visual Content)

The Synthetic SVC (SSVC) dataset comprises 12,000 images with respective bounding box annotations and detailed graph representations. This dataset enables the development of strong models for the interpretation of SVCs while skipping the time-consuming dense data annotation.

1 papers0 benchmarks

Fire and Smoke Dataset

This dataset is collected by DataCluster Labs, India. To download full dataset or to submit a request for your new data collection needs, please drop a mail to: sales@datacluster.ai

1 papers0 benchmarksImages

Nakdimon-train

A collection of diacritized Hebrew text in a variety of registers and from different sources.

1 papers0 benchmarksTexts

Natural sentences that contain *any*

We scraped the Gutenberg Project and a subset of English Wikipedia to obtain the list of sentences that contain any. Next, using a combination of heuristics, we filtered the result with regular expressions to produce two sets of sentences (the second set underwent additional manual filtration): * 3844 sentences with sentential negation and a plural object with any to the right to the verb; * 330 sentences with nobody / no one as subject and a plural object with any to the right.

1 papers0 benchmarksTexts

Synthetic parallel sentences that contain *any*

We used the following procedure. First, we automatically identified the set of verbs and nouns to build our items from. To do so, we started with bert-base-uncased vocabulary. We ran all non-subword lexical tokens through a SpaCy POS. Further, we lemmatized the result using https://pypi.org/project/Pattern/ and dropped duplicates. Then, we filtered out modal verbs, singularia tantum nouns and some visible lemmatization mistakes. Finally, we filtered out non-transitive verbs to give the dataset a bit of a higher baseline of grammaticality.

1 papers0 benchmarksTexts

Simulated micro-Doppler Signatures

Simulated pulse Doppler radar signatures for four classes of helicopter-like targets. The classes differ in the number of rotating blades each kind of target carries, thus each class translates into a specific modulation pattern on the Doppler signature. Doppler signatures are a typical feature used to achieve radar targets discrimination. This dataset was generated using a simple open-source MATLAB simulation code, which can be easily modified to generate custom datasets with more classes and increased intra-class diversity.

1 papers0 benchmarks

Extended Minecraft Corpus dataset

Minecraft Corpus dataset with builder utterance annotations

1 papers0 benchmarksImages, Texts

Replication Data for: Investigating the concentration of High Yield Investment Programs in the United Kingdom

The dataset provides information about 450 HYIPs collected between November 2020 and September 2021. This dataset was analyzed and the results are discussed in the paper.

1 papers0 benchmarksTables

Twitter MediaEval (MediaEval Benchmarking Initiative for Multimedia Evaluation)

The task addresses the problem of the appearance and propagation of posts that share misleading multimedia content (images or video). In the context of the task, different types of misleading use are considered:

1 papers0 benchmarksImages, Texts

WikiBanEvasion (Wikipedia Ban Evasion Dataset)

A dataset comprising 8,551 ban evasion pairs on Wikipedia, where each pair comprises a parent account and the child account. We adopt a strategy to ensure that there is a 1:1 mapping between parent and child accounts. For each of the accounts in these ban evasion pairs, we provide the following data: - Wikipedia usernames, creation date, ban date, and other account-level meta-data - Corresponding edit information in form of revision IDs, pages edited, added text, deleted text, edit comment, and timestamp

1 papers0 benchmarksTexts

Water Footprint Recommender System Data

It contains data from two different realities: Food.com, a well-known American recipe site, and Planeat, an Italian site that allows you to plan recipes to save food waste. The dataset is divided into two parts: embeddings, which can be used directly to execute the work and receive suggestions, and raw data, which must first be processed into embeddings.

1 papers0 benchmarksTables, Texts

OAGT (Paper Topic Dataset)

OAGL is a paper topic dataset consisting of 6942930 records which comprise various scientific publication attributes like abstracts, titles, keywords, publication years, venues, etc. The last two fields of each record are the topic id from a taxonomy of 27 topics created from the entire collection and the 20 most significant topic words. Each dataset record (sample) is stored as a JSON line in the text file.

1 papers0 benchmarksTexts

Bosch CNC Machining Dataset

The dataset provided is a collection of real-world industrial vibration data collected from a brownfield CNC milling machine. The acceleration has been measured using a tri-axial accelerometer (Bosch CISS Sensor) mounted inside the machine. The X- Y- and Z-axes of the accelerometer have been recorded using a sampling rate equal to 2 kHz. Thereby normal as well as anomalous data have been collected for 4 different timeframes, each lasting 5 months from February 2019 until August 2021 and labelled accordingly. It can be used to investigate the scalability of models and research process variations as the anomaly impact differs. In total there is data from three different CNC milling machines each executing 15 processes. For a detailed description of the data and experimental set-up, please refer to the paper: https://doi.org/10.1016/j.procir.2022.04.022

1 papers0 benchmarksTime series

Heroes Corpus

Each episode directory contains word-level and segment-level information of the whole episode and also parallel samples extracted under segments_eng and segments_spa subdirectories. Each sample is stored as an WAV audio file, text file and a CSV file containing word timing information and word-level paralinguistic and prosodic features.

1 papers0 benchmarks
PreviousPage 429 of 1000Next