19,997 machine learning datasets
19,997 dataset results
The SuperLim dataset is a Swedish version of the English benchmarking platform (Super)GLUE. It forms the basis of a national testbed for Swedish language models. The project aims to provide a standardized collection of benchmarking tests for Swedish language models, supporting the development of trustworthy and robust Natural Language Processing (NLP) applications.
The DanishPoliticalComments dataset is a collection of sentences that are labeled with fine-grained polarity in the range from -2 to 2 (negative to positive). The dataset is primarily used for tasks related to text classification, specifically multi-class classification. The dataset is monolingual and is in the Danish language. It falls under the size category of 1K<n<10K. The language creators are listed as 'other' and the annotations are expert-generated.
This dataset is suitable for sentiment analysis, which consists of Danish data from the Leipzig Collection. The collection and annotation of the dataset are solely due to Finn Årup Nielsen. It was originally annotated as a score between -5 and +5, but the labels in this version have been converted to negative, neutral, and positive labels.
bioRxiv is a free online archive for unpublished preprints in the life sciences. It allows researchers to share their findings with the scientific community and receive feedback before it undergoes peer review for formal publication. Datasets on bioRxiv are typically associated with these preprints.
This dataset can be used as a benchmark for clustering word embeddings for German. It contains 18'084 unique samples, 28 splits with 177 to 16'425 samples, and 4 to 93 unique classes.
medRxiv (pronounced "med-archive") is an Internet site distributing unpublished e-prints about health sciences. It distributes complete but unpublished manuscripts in the areas of medicine, clinical research, and related health sciences without charge to the reader. Such manuscripts have yet to undergo peer review and the site notes that preliminary status and that the manuscripts should not be considered for clinical application, nor relied upon for news reporting as established information.
Given two sentences, the participants are asked to determine whether they express the same or very similar meaning and optionally a degree score between 0 and 1. Following the literature on paraphrase identification, we evaluate system performance primarily by the F-1 score and Accuracy against human judgments. We also provide additional evaluations by Pearson correlation and PINC (Chen and Dolan, 2011), which measure lexical dissimilarity between sentence pairs.
The LinkSO dataset is a resource for learning to retrieve similar question-answer pairs on Stack Overflow. It consists of three datasets corresponding to three popular programming languages (Python, Java, JavaScript), 690K question pairs, and 26K linked question pairs (i.e., positive examples). The dataset was extracted from Stack Overflow's data dump in April 2018 and was cleaned and pre-processed to remove non-ASCII characters, email addresses, URLs, and code blocks. The dataset was designed to help propose new models, such as neural network models, to improve community-based question-answer retrieval in the software engineering domain.
The SweFAQ dataset is a collection of frequently asked questions from Swedish authorities' websites with shuffled answers. It was created by Aleksandrs Berdicevskis and is published by Språkbanken Text. The dataset is a part of the SuperLim collection.
The cMedQA dataset is designed for Chinese community medical question answering. It has two versions:
The QQ Browser Query Title Corpus (QBQTC) is a large-scale dataset constructed for search scenarios by the QQ Browser search engine. It integrates dimensions such as relevance, authority, content quality, and timeliness, and is widely used in search engine business scenarios.
A collection of multilingual sentiment datasets grouped into 3 classes -- positive, neutral, and negative.
The Stockholm-Umeå Corpus (SUC) is a collection of Swedish texts from the 1990s, consisting of one million words in total. The corpus is balanced, meaning that it contains various text types and stylistic levels. The texts are annotated with part-of-speech tags, morphological analysis, and lemma (all that can be considered gold standard data), as well as some structural and functional information.
WMT21 (Workshop on Machine Translation 2021) Translation Task focuses on news text translation. It includes language pairs such as English to/from various languages like Chinese, Czech, German, Hausa, Icelandic, Japanese, Russian, and more. Goals include investigating current MT techniques for languages other than English, challenges in translating between language families, translation of low-resource languages, and creating publicly available corpora for MT evaluation. The dataset provides parallel corpora for all languages and additional resources for download, with a focus on machine translation of news.
The TICO-19 dataset is a translation initiative focused on COVID-19 content, created by academic and industry partners along with Translators without Borders. It includes translation memories, translated terminologies for COVID-19-related terms, and a benchmark dataset. The benchmark comprises 30 documents, translating 3071 sentences (69.7k words) from English into 36 languages. The effort aims to assist professional translators and machine translation research, emphasizing emergency and crisis-related content availability in multiple languages.
Dataset for range data gathered using the Ultra-wide Band (UWB) MDEK 1001 Dev. kit in 3 different experimental scenarios. The dataset description document provides the details of experimentation and the dataset.
This dataset provides daily weather information for capital cities around the world. Unlike forecast data, this dataset offers a comprehensive set of features that reflect the current weather conditions around the world. Starting from August 29, 2023. It provides over 40+ features , including temperature, wind, pressure, precipitation, humidity, visibility, air quality measurements and more. The dataset is valuable for analyzing Global weather patterns, exploring climate trends, and understanding the relationships between different weather parameters.
A dataset specifically tailored to the biotech news sector, aiming to transcend the limitations of existing benchmarks. This dataset is rich in complex content, comprising various biotech news articles covering various events, thus providing a more nuanced view of information extraction challenges.
Technologically Assisted Reviews in Empirical Medicine.
This dataset comprises micro-ultrasound scans and human prostate annotations of 75 patients who underwent micro-ultrasound guided prostate biopsy at the University of Florida. All images and segmentations have been fully de-identified in the NIFTI format.