19,997 machine learning datasets
19,997 dataset results
India is a linguistic area with one of the longest histories of contact, influence, use, teaching and learning of English-in-diaspora in the world (Kachru and Nelson, 2006). Thus, a huge number of Indians active on the internet are able in English communication to some degree. India also enjoys huge diversity in language. Apart from Hindi, it has several regional languages that are the primary tongue of people native to the region. This is to the extent that social media including Facebook, WhatsApp, Twitter, etc. contain more than one language, and such phenomena are called code-mixing and code-switching. On the other side, the evolution of sentiments from such social media texts have also created many new opportunities for information access and language technology, but also many new challenges, making it one of the prime present-day research areas. Sentiment analysis in code-mixed data has several real-life applications in opinion mining from social media campaign to feedback analys
Context The Kepler Space Observatory is a NASA-build satellite that was launched in 2009. The telescope is dedicated to searching for exoplanets in star systems besides our own, with the ultimate goal of possibly finding other habitable planets besides our own. The original mission ended in 2013 due to mechanical failures, but the telescope has nevertheless been functional since 2014 on a "K2" extended mission.
Reasoning over spans of tokens from different parts of the input is essential for natural language understanding (NLU) tasks such as fact-checking (FC), machine reading comprehension (MRC) or natural language inference (NLI). We introduce SpanEx, a multi-annotator dataset of human-annotated span interaction explanations for two NLU tasks: NLI and FC.
A multimodal dataset of radio galaxies and their corresponding infrared hosts.
Neural fields (NeFs) have recently emerged as a versatile method for modeling signals of various modalities, including images, shapes, and scenes. Subsequently, many works have explored the use of NeFs as representations for downstream tasks, e.g. classifying an image based on the parameters of a NeF that has been fit to it. However, the impact of the NeF hyperparameters on their quality as downstream representation is scarcely understood and remains largely unexplored. This is partly caused by the large amount of time required to fit datasets of neural fields.
Hand-labelled dataset of crop and non-crop labels distributed throughout Nigeria with respective hd5f data arrays.
Pre-training dataset used in paper "From Artificially Real to Real: Leveraging Pseudo Data from Large Language Models for Low-Resource Molecule Discovery" (AAAI 2024)
This is a cleaned version of the dataset introduced in Kaggle by user SAJID576. The original dataset's link is https://www.kaggle.com/datasets/sajid576/sql-injection-dataset
The Generic Object Decoding (GOD) Dataset is a specialized resource developed for fMRI-based decoding. It aggregates fMRI data gathered through the presentation of images from 200 representative object categories, originating from the 2011 fall release of ImageNet. The training session incorporated 1,200 images (8 per category from 150 distinct object categories). In contrast, the test session included 50 images (one from each of the 50 object categories). It is noteworthy that the categories in the test session were unique from those in the training session and were introduced in a randomized sequence across runs. On five subjects the fMRI scanning was conducted.
CCGR (Cross-Covariate Gait Recognition), the first gait dataset for studying cross-covariate challenges, contains 970 subjects, approximately 1.6 million sequences, 53 covariates, and 33 views.
Aircraft collision avoidance systems rely on sensor information to detect and track intruding aircraft so that they may issue proper collision avoidance advisories. While typical surveillance sensors for manned aircraft include transponders and onboard radar, autonomous aircraft will require additional sensors both for redundancy and to replace the visual acquisition typically performed by the pilot. As a result, the community has proposed detecting other aircraft using vision-based sensors such as cameras. These sensors require the development of techniques to process images of the environment to detect intruding aircraft. To boost this development, this artifact provides a dataset of 72,000 labeled images of intruder aircraft with various lighting conditions, weather conditions, relative geometries, and geographic locations. For more information on the structure of this dataset as well as benchmark models and a full simulator, see https://github.com/sisl/VisionBasedAircraftDAA.
This repository is an extension of GEval. This repository contains a (software) evaluation framework to perform evaluation and comparison on RDF-star graph embedding techniques. The gold standard datasets for evaluation were created from KGRC-RDF-star. Please see here.
$\textbf{Turbulence}$ is a new benchmark for systematically evaluating the correctness and robustness of instruction-tuned large language models (LLMs) for code generation. Turbulence consists of a large set of natural language question templates, each of which is a programming problem, parameterised so that it can be asked in many different forms. Each question template has an associated test oracle that judges whether a code solution returned by an LLM is correct. Thus, from a single question template, it is possible to ask an LLM a $\textit{neighbourhood}$ of very similar programming questions, and assess the correctness of the result returned for each question. This new benchmark systematically and automatically identifies cases where LLMs are able to solve some problems in a neighbourhood but do not manage to generalise to solve the whole neighbourhood. Therefore, this method is effective at highlighting robustness issues.
Official dataset for Towards A Holistic Landscape of Situated Theory of Mind in Large Language Models. Ziqiao Ma, Jacob Sansom, Run Peng, Joyce Chai. EMNLP Findings, 2023.
The LTI LangID Corpus is a dataset used for language identification (LangID) tasks. It contains text data in various languages. The dataset has had multiple releases, with the first release containing 781 "core" languages and 1091 languages overall.
The Wikimedia dataset refers to a collection of data related to Wikimedia projects, which include Wikipedia, Wikivoyage, Wiktionary, Wikisource, and others. These datasets are publicly available and can be used for various purposes such as research, backup, and offline use.
The MegaAcceptability dataset is a collection of ordinal acceptability judgments for clause-embedding verbs of English in various surface-syntactic frames and matrix tenses. Here are some key details: - It includes judgments for 1,007 clause-embedding verbs of English in 50 surface-syntactic frames and three matrix tenses. - The dataset combines the MegaAcceptability version 1.0 and data collected for 25,000 additional verb-frame pairs on Amazon’s Mechanical Turk using Ibex on Mechanical Turk. - The dataset has been used to address questions in linguistic theory.
See https://zenodo.org/records/8340383
See https://zenodo.org/records/10185500
The Amazon Polarity dataset is a set of reviews from Amazon. The dataset is constructed by taking review scores 1 and 2 as negative (class 1), and 4 and 5 as positive (class 2). Reviews with a score of 3 are ignored. The dataset spans a period of 18 years, including approximately 35 million reviews up to March 2013. Each class in the dataset has 1,800,000 training samples and 200,000 testing samples. The dataset includes product and user information, ratings, and a plaintext review.