TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Second HAREM (Segundo HAREM)

The Second HAREM was an evaluation exercise in Portuguese Named Entity Recognition. It aims to refine text annotation processes, building on the First HAREM. Challenges include adapting guidelines for new texts and establishing a unified document with directives from both editions.

0 papers0 benchmarksTexts

SIGARRA News Corpus

This dataset was taken from the SIGARRA information system at the University of Porto (UP). Every organic unit has its own domain and produces academic news. We collected a sample of 1000 news, manually annotating 905 using the Brat rapid annotation tool. This dataset consists of three files. The first is a CSV file containing news published between 2016-12-14 and 2017-03-01. The second file is a ZIP archive containing one directory per organic unit, with a text file and an annotations file per news article. The third file is an XML containing the complete set of news in a similar format to the HAREM dataset format. This dataset is particularly adequate for training named entity recognition models.

0 papers0 benchmarksTexts

PropBank-PT

The PropBankPT (Branco et al., 2012) is a set of sentences annotated with their constituency structure and semantic role tags, composed of 3,406 sentences and 44,598 tokens taken from the Wall Street Journal translated. For the creation of this PropBank we adopted a semi-automatic analysis with a double-blind annotation followed by adjudication. The resulting dataset contains three information levels: phrase constituency, grammatical functions, and phrase semantic roles. The main motivation behind the creation of this resource was to build a high quality data set with semantic information that could support the development of automatic semantic role labelers for Portuguese. The development of this resource started under the METANET4U project (at: http://metanet4u.eu/) whose main goal is to contribute to the establishment of a pan-European digital platform that makes available language resources and services, encompassing both datasets and software tools, for speech and language process

0 papers0 benchmarksTexts

Mac-Morpho

Mac-Morpho is a corpus of Brazilian Portuguese texts annotated with part-of-speech tags. Its first version was released in 2003 [1], and since then, two revisions have been made in order to improve the quality of the resource [2, 3]. The corpus is available for download split into train, development and test sections. These are 76%, 4% and 20% of the corpus total, respectively (the reason for the unusual numbers is that the corpus was first split into 80%/20% train/test, and then 5% of the train section was set aside for development). This split was used in [3], and new POS tagging research with Mac-Morpho is encouraged to follow it in order to make consistent comparisons possible.

0 papers0 benchmarksTexts

Benchmarking-Chinese-Text-Recognition

This repository contains datasets and baselines for benchmarking Chinese text recognition. Please see the corresponding paper for more details regarding the datasets, baselines, the empirical study, etc.

0 papers0 benchmarks

RuSRL

This dataset contains annotations of semantic frames and intra-frame syntax for 1500 Russian sentences. Each sentence is annotated with predicate-argument structures. Syntactic information is also provided for each frame.

0 papers0 benchmarks

Blood Metabolomic and Preeclampsia A Mendelian Randomization Study (Ling Keng)

Supplemental Rcode with original results and images

0 papers0 benchmarks

PhyMER (PhyMER: Physiological Dataset for Multimodal Emotion Recognition With Personality as a Context)

A physiological signal dataset with multiple physiological signals collected from 30 participants. The dataset consists of EEG, EDA, BVP, and temperature information with annotations in both categorical and dimensional views along with the individual personality traits of the participants for the study of emotions in presence of individual personality differences. The PhyMER dataset consists of the recorded physiological signals, participants' personality information, and emotion annotations in terms of Arousal, Valence, and seven basic emotions.

0 papers0 benchmarksEEG

Digit Sign

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

0 papers0 benchmarks

EyePACS-light (v2) (EyePACS-AIROGS-light-v2)

This is an improved machine-learning-ready glaucoma dataset using a balanced subset of standardized fundus images from the Rotterdam EyePACS AIROGS [1] set. This dataset is split into training, validation, and test folders which contain 4000 (~84%), 385 (~8%), and 385 (~8%) fundus images in each class respectively. Each training set has a folder for each class: referable glaucoma (RG) and non-referable glaucoma (NRG).

0 papers0 benchmarksImages, Medical

HouseCat6D (A Large-Scale Multi-Modal Category Level 6D Object Perception Dataset with Household Objects in Realistic Scenarios)

Estimating 6D object poses is a major challenge in 3D computer vision. Building on successful instance-level approaches, research is shifting towards category-level pose estimation for practical applications. Current categorylevel datasets, however, fall short in annotation quality and pose variety. Addressing this, we introduce HouseCat6D, a new category-level 6D pose dataset. It features 1) multimodality with Polarimetric RGB and Depth (RGBD+P), 2) encompasses 194 diverse objects across 10 household categories, including two photometrically challenging ones, and 3) provides high-quality pose annotations with an error range of only 1.35 mm to 1.74 mm. The dataset also includes 4) 41 large-scale scenes with comprehensive viewpoint and occlusion coverage, 5) a checkerboard-free environment, and 6) dense 6D parallel-jaw robotic grasp annotations. Additionally, we present benchmark results for leading category-level pose estimation networks.

0 papers0 benchmarksRGB-D

I-CARE: International Cardiac Arrest REsearch consortium Database

The International Cardiac Arrest REsearch consortium (I-CARE) Database includes baseline clinical information and continuous electroencephalogram (EEG) and electrocardiogram (ECG) recordings from comatose patients following cardiac arrest. The patients were admitted to an intensive care unit (ICU) in one of seven academic hospitals in the U.S. and Europe and monitored for several hours to several days. The long-term neurological function of the patients was determined using the Cerebral Performance Category scale.

0 papers0 benchmarksEEG

Matrix Shapes

The task aims to measure the capability of models to predict the shape of the result of a chain of matrix manipulations, given the inputs' shapes. This involves knowledge of the effect of individual manipulations as well as the ability to combine this knowledge (multi-hop inference).

0 papers0 benchmarks

Data Wrangling

This dataset is part of the Data Wrangling Dataset Repository created by the DMiP Team (UPV). The term "data wrangling" usually refers to a great deal of repetitive and very time-consuming data preparation tasks, such as the acquisition, integration, manipulation, cleansing, enriching, and transformation of data. All the datasets include six examples of one particular problem, with an input and the expected output.

0 papers0 benchmarks

MegaIntensionality

The MegaIntensionality dataset is a part of the MegaAttitude project. It consists of slider-based judgments of doxastic and balletic inferences for 725 finite clause-embedding verbs of English with a variety of subordinate clause structures, matrix tenses, and matrix subjects. The dataset is used to study lexically triggered inferences related to predicates' intensional properties, particularly those related to the belief or desire of one or both participants. These inferences are of interest due to the patterns that emerge in how the inferences are affected by contexts such as negation.

0 papers0 benchmarks

MegaOrientation

The MegaOrientation Dataset is a linguistic resource that consists of ordinal acceptability judgments for 898 clause-embedding verbs of English with a variety of nonfinite subordinate clause structures. This dataset is used to examine aspects of semantic interpretation that are due to predicates’ denotations and those that are due to the denotations of their arguments. It particularly focuses on the context of temporal interpretation, which acts as an indicator of underlying syntactic structures and semantic frames.

0 papers0 benchmarks

COCA (The Corpus of Contemporary American English)

The Corpus of Contemporary American English (COCA) is a large and balanced corpus of American English. It contains more than one billion words of text (25+ million words each year from 1990 to 2019) from eight genres: spoken, fiction, popular magazines, newspapers, academic texts, TV and Movie subtitles, blogs, and other web pages. COCA is probably the most widely-used corpus of English and it offers unparalleled insight into variation in English.

0 papers0 benchmarks

Wikidata

Wikidata is a free and open knowledge base that can be read and edited by both humans and machines. It acts as central storage for the structured data of its Wikimedia sister projects including Wikipedia, Wikivoyage, Wiktionary, Wikisource, and others.

0 papers0 benchmarks

SIT (Symbol Interpretation Task)

This task is composed of five different subtasks that require interpreting statements referring to structures of a simple world. This world is built using emojis; a structure of the world is simply a sequence of six emojis. Crucially, in every variation, we make explicit the semantic link between the emojis and their name in a different way:

0 papers0 benchmarks

Trillion Word Corpus (10,000 most common English words)

The Trillion Word Corpus is a dataset created by Google, which contains one trillion words from public web pages. It was developed by harnessing the vast power of Google's data centers and distributed processing infrastructure to process larger and larger training corpora.

0 papers0 benchmarks
PreviousPage 655 of 1000Next