19,997 machine learning datasets
19,997 dataset results
The file contains an annotated list of papers that are included in the literature survey.
BioFuelQR is a dataset consisting of complex reasoning questions related to catalyst discovery in biofuels. This dataset is aimed at benchmarking scientific question answering methods, particularly for search based text generation.
DotPrompts is a set of testcases derived from PragmaticCode, such that each testcase consists of a prompt to a dereference location (a code location having the "." operator in Java). It is primarily meant as a benchmark for Code LMs.
A mapping of Quasimodo to the relations of ConceptNet.
A mapping of Ascent to the relations of ConceptNet.
We present the SourceData-NLP dataset produced through the routine curation of papers during the publication process. A unique feature of this dataset is its emphasis on the annotation of bioentities in figure legends. We annotate eight classes of biomedical entities (small molecules, gene products, subcellular components, cell lines, cell types, tissues, organisms, and diseases), their role in the experimental design, and the nature of the experimental method as an additional class. SourceData-NLP contains more than 620,000 annotated biomedical entities, curated from 18,689 figures in 3,223 papers in molecular and cell biology. We illustrate the dataset's usefulness by assessing BioLinkBERT and PubmedBERT, two transformers-based models, fine-tuned on the SourceData-NLP dataset for NER. We also introduce a novel context-dependent semantic task that infers whether an entity is the target of a controlled intervention or the object of measurement.
Reader eye tracking and engagement scores for two short stories, aggregated by sentence.
A fully-annotated, open-design dataset of autonomous and piloted high-speed flight
Test set of sentences in Hindi with complex coreference involving two entities inspired by WinoBias format of sentences in English. Includes grammatical gender cues of Hindi to test gender bias in Hindi-English NMT Systems.
Test set of sentences in Hindi with simple gender-specific context used to measure gender bias in NMT systems for Hindi-English.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
This dataset was built with data acquired at the Hospital Clinic of Barcelona, Spain. It is composed of a total of 1126 HD polyp images. There are a total of 473 unique polyps, with a variable number of different shots per polyp (minimum: 2, maximum: 24, median: 10). Special attention was paid to ensure that images from the same polyp show different conditions. An external frame-grabber and a white light endoscope were used to capture raw images. The dataset contains images with two different resolutions: 1920 x 1080 and 1350 x 1080.
The FoodSG-233 dataset contains 209,861 images, covering 13 food groups and 233 food categories.
This dataset contains network traces collected in-lab and in a real-world setting. We also collected ground truth Quality of Experience (QoE) logs from the browser. Researchers can use this dataset to solve a wide range of problems that involve understanding video conferencing quality from a network perspective.
A dataset of natural language data collected by putting together more than 150 existing mono-lingual and multilingual datasets together and crawling known multilingual websites. The focus of this dataset is on 500 extremely low-resource languages.
Unsynchronized dynamic blender dataset for multi-view dynamic NeRFs for evaluating MAE between the predicted time offsets and the ground truth. It contains three unsyncrhonized scenes; box, deer, and fox.
Versatile synthetic classification dataset based on precise input spike timings drawn from smooth random manifolds as previously described
Faces Through Time (FTT) features 26,247 images of notable people from the 19th to 21st centuries, with roughly 1,900 images per decade on average. It is sourced from Wikimedia Commons, a crowdsourced and open-licensed collection of 50M images.
we have prepared a dataset using publicly available TED Talks transcripts [27] and selected the Turkish corpus. The resulting Turkish punctuation restoration dataset currently consists of 146K sentences and 1.8M tokens. The ratio of the train, validation, and test splits are 0.8, 0.1, and 0.1, respectively. Data files contain two columns. The first column has the tokens separated by white space. The second column includes tags for each token.