19,997 machine learning datasets
19,997 dataset results
The Second HAREM was an evaluation exercise in Portuguese Named Entity Recognition. It aims to refine text annotation processes, building on the First HAREM. Challenges include adapting guidelines for new texts and establishing a unified document with directives from both editions.
This dataset was taken from the SIGARRA information system at the University of Porto (UP). Every organic unit has its own domain and produces academic news. We collected a sample of 1000 news, manually annotating 905 using the Brat rapid annotation tool. This dataset consists of three files. The first is a CSV file containing news published between 2016-12-14 and 2017-03-01. The second file is a ZIP archive containing one directory per organic unit, with a text file and an annotations file per news article. The third file is an XML containing the complete set of news in a similar format to the HAREM dataset format. This dataset is particularly adequate for training named entity recognition models.
The PropBankPT (Branco et al., 2012) is a set of sentences annotated with their constituency structure and semantic role tags, composed of 3,406 sentences and 44,598 tokens taken from the Wall Street Journal translated. For the creation of this PropBank we adopted a semi-automatic analysis with a double-blind annotation followed by adjudication. The resulting dataset contains three information levels: phrase constituency, grammatical functions, and phrase semantic roles. The main motivation behind the creation of this resource was to build a high quality data set with semantic information that could support the development of automatic semantic role labelers for Portuguese. The development of this resource started under the METANET4U project (at: http://metanet4u.eu/) whose main goal is to contribute to the establishment of a pan-European digital platform that makes available language resources and services, encompassing both datasets and software tools, for speech and language process
Mac-Morpho is a corpus of Brazilian Portuguese texts annotated with part-of-speech tags. Its first version was released in 2003 [1], and since then, two revisions have been made in order to improve the quality of the resource [2, 3]. The corpus is available for download split into train, development and test sections. These are 76%, 4% and 20% of the corpus total, respectively (the reason for the unusual numbers is that the corpus was first split into 80%/20% train/test, and then 5% of the train section was set aside for development). This split was used in [3], and new POS tagging research with Mac-Morpho is encouraged to follow it in order to make consistent comparisons possible.
This repository contains datasets and baselines for benchmarking Chinese text recognition. Please see the corresponding paper for more details regarding the datasets, baselines, the empirical study, etc.
This dataset contains annotations of semantic frames and intra-frame syntax for 1500 Russian sentences. Each sentence is annotated with predicate-argument structures. Syntactic information is also provided for each frame.
Supplemental Rcode with original results and images
A physiological signal dataset with multiple physiological signals collected from 30 participants. The dataset consists of EEG, EDA, BVP, and temperature information with annotations in both categorical and dimensional views along with the individual personality traits of the participants for the study of emotions in presence of individual personality differences. The PhyMER dataset consists of the recorded physiological signals, participants' personality information, and emotion annotations in terms of Arousal, Valence, and seven basic emotions.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
This is an improved machine-learning-ready glaucoma dataset using a balanced subset of standardized fundus images from the Rotterdam EyePACS AIROGS [1] set. This dataset is split into training, validation, and test folders which contain 4000 (~84%), 385 (~8%), and 385 (~8%) fundus images in each class respectively. Each training set has a folder for each class: referable glaucoma (RG) and non-referable glaucoma (NRG).
Estimating 6D object poses is a major challenge in 3D computer vision. Building on successful instance-level approaches, research is shifting towards category-level pose estimation for practical applications. Current categorylevel datasets, however, fall short in annotation quality and pose variety. Addressing this, we introduce HouseCat6D, a new category-level 6D pose dataset. It features 1) multimodality with Polarimetric RGB and Depth (RGBD+P), 2) encompasses 194 diverse objects across 10 household categories, including two photometrically challenging ones, and 3) provides high-quality pose annotations with an error range of only 1.35 mm to 1.74 mm. The dataset also includes 4) 41 large-scale scenes with comprehensive viewpoint and occlusion coverage, 5) a checkerboard-free environment, and 6) dense 6D parallel-jaw robotic grasp annotations. Additionally, we present benchmark results for leading category-level pose estimation networks.
The International Cardiac Arrest REsearch consortium (I-CARE) Database includes baseline clinical information and continuous electroencephalogram (EEG) and electrocardiogram (ECG) recordings from comatose patients following cardiac arrest. The patients were admitted to an intensive care unit (ICU) in one of seven academic hospitals in the U.S. and Europe and monitored for several hours to several days. The long-term neurological function of the patients was determined using the Cerebral Performance Category scale.
The task aims to measure the capability of models to predict the shape of the result of a chain of matrix manipulations, given the inputs' shapes. This involves knowledge of the effect of individual manipulations as well as the ability to combine this knowledge (multi-hop inference).
This dataset is part of the Data Wrangling Dataset Repository created by the DMiP Team (UPV). The term "data wrangling" usually refers to a great deal of repetitive and very time-consuming data preparation tasks, such as the acquisition, integration, manipulation, cleansing, enriching, and transformation of data. All the datasets include six examples of one particular problem, with an input and the expected output.
The MegaIntensionality dataset is a part of the MegaAttitude project. It consists of slider-based judgments of doxastic and balletic inferences for 725 finite clause-embedding verbs of English with a variety of subordinate clause structures, matrix tenses, and matrix subjects. The dataset is used to study lexically triggered inferences related to predicates' intensional properties, particularly those related to the belief or desire of one or both participants. These inferences are of interest due to the patterns that emerge in how the inferences are affected by contexts such as negation.
The MegaOrientation Dataset is a linguistic resource that consists of ordinal acceptability judgments for 898 clause-embedding verbs of English with a variety of nonfinite subordinate clause structures. This dataset is used to examine aspects of semantic interpretation that are due to predicates’ denotations and those that are due to the denotations of their arguments. It particularly focuses on the context of temporal interpretation, which acts as an indicator of underlying syntactic structures and semantic frames.
The Corpus of Contemporary American English (COCA) is a large and balanced corpus of American English. It contains more than one billion words of text (25+ million words each year from 1990 to 2019) from eight genres: spoken, fiction, popular magazines, newspapers, academic texts, TV and Movie subtitles, blogs, and other web pages. COCA is probably the most widely-used corpus of English and it offers unparalleled insight into variation in English.
Wikidata is a free and open knowledge base that can be read and edited by both humans and machines. It acts as central storage for the structured data of its Wikimedia sister projects including Wikipedia, Wikivoyage, Wiktionary, Wikisource, and others.
This task is composed of five different subtasks that require interpreting statements referring to structures of a simple world. This world is built using emojis; a structure of the world is simply a sequence of six emojis. Crucially, in every variation, we make explicit the semantic link between the emojis and their name in a different way:
The Trillion Word Corpus is a dataset created by Google, which contains one trillion words from public web pages. It was developed by harnessing the vast power of Google's data centers and distributed processing infrastructure to process larger and larger training corpora.