19,997 machine learning datasets
19,997 dataset results
The Angry Tweets dataset is a collection of anonymized Danish Twitter data that has been annotated for sentiment analysis through crowd-sourcing. Here are some key details about the dataset:
The 10kGNAD dataset is intended to solve part of this problem as the first German topic classification dataset. It consists of 10273 German-language news articles from an Austrian online newspaper categorized into nine topics. These articles are a now unused part of the One Million Posts Corpus.
A Crowdsourced Multi-Domain Music Dataset of Europeana Collections
The Touché23-ValueEval Dataset is a collection of arguments used for identifying human values behind those arguments. It was created by collecting 9324 arguments from 6 diverse sources, including religious texts, political discussions, free-text arguments, newspaper editorials, and online democracy platforms. Each argument was annotated by 3 crowdworkers for 54 values.
The cCOVID-News dataset is a publicly available Chinese text retrieval dataset created from COVID-19 news articles. It contains a collection of text data related to COVID-19, and it is used as part of the out-of-domain evaluation for the DuReader retrieval benchmark.
The Multilingual Low-Resource Translation task for Indo-European Languages, part of the EMNLP 2021 Conference, focused on improving machine translation in the cultural heritage domain for North-Germanic and Romance languages. It aimed to explore data transferability across related languages, prioritizing low-resource languages while allowing training in high-resource languages. The task had two subtasks: translating Europeana thesis abstracts and descriptions for North-Germanic languages, and translating Wikipedia cultural heritage articles for Romance languages. The task encouraged using diverse data sources and provided additional resources like lexicons and validation sets. Evaluation was based on translation quality, emphasizing multilinguality and resource efficiency in machine translation.
TACO (Topics in Algorithmic Code generation dataset) is a dataset focused on algorithmic code generation, designed to provide a more challenging training dataset and evaluation benchmark for the code generation model field. The dataset consists of programming competition problems that are more difficult and closer to real programming scenarios. It emphasizes improving or evaluating the model's understanding and reasoning abilities in practical application scenarios, rather than just implementing predefined function functionalities.
The Tiny Shakespeare corpus is a dataset that contains 40,000 lines of Shakespeare from a variety of his plays. The Tiny Shakespeare corpus is a popular choice for training language models due to its manageable size and the complexity of Shakespeare's language. It provides a good balance between computational efficiency and the ability to generate interesting text.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
TuPyE, an enhanced iteration of TuPy, encompasses a compilation of 43,668 meticulously annotated documents specifically selected for the purpose of hate speech detection within diverse social network contexts. This augmented dataset integrates supplementary annotations and amalgamates with datasets sourced from Fortuna et al. (2019), Leite et al. (2020), and Vargas et al. (2022), complemented by an infusion of 10,000 original documents from the TuPy-Dataset.
KAgentBench is a benchmark dataset of over 3,000 human-edited, automated evaluation data for testing agent capabilities, with evaluation dimensions including planning, tool use, reflection, concluding, and profiling.
Latent DNA Diffusion Dataset
This is the sequential instructions dataset from Understanding the Effects of RLHF on LLM Generalisation and Diversity. The dataset is in the alpaca_eval format.
ODSI-DB is an image database of oral and dental reflectance spectral images of human test subjects. Image sets of the test subjects contain the front-view and the occlusal surfaces of lower and upper teeth, oral mucosa, and face surrounding the mouth. Other features-of-interest have been imaged on case-by-case basis. The spectral images in the database have been annotated by dental experts.
Click to #Development of QA-SQP for non-linear and history-dependent mechanical problems
Dataset introduction There are four dimension in MBTI. And there are two opposite attributes within each dimension.
ABSTRACT Recently, the technology of the fourth revolution has given the characteristics of things constantly expanding, and everything, including people, things, people, and the environment, is connected based on the Internet. In particular, the network structure is connected to various IoT devices and is changing from wired to wireless. Unlike users who operated each device, other devices can now be operated through gateways inside and outside the smart home. However, these changes have created an environment vulnerable to external attacks, and when an attacker accesses a gateway, he can attempt various attacks, including Port scans, OS&Service detection, and DoS attacks on IoT devices. Therefore, we disclose the dataset below to promote security research on IoT.
This dataset contains news headlines relevant to key forex pairs: AUDUSD, EURCHF, EURUSD, GBPUSD, and USDJPY. The data was extracted from reputable platforms Forex Live and FXstreet over a period of 86 days, from January to May 2023. The dataset comprises 2,291 unique news headlines. Each headline includes an associated forex pair, timestamp, source, author, URL, and the corresponding article text. Data was collected using web scraping techniques executed via a custom service on a virtual machine. This service periodically retrieves the latest news for a specified forex pair (ticker) from each platform, parsing all available information. The collected data is then processed to extract details such as the article's timestamp, author, and URL. The URL is further used to retrieve the full text of each article. This data acquisition process repeats approximately every 15 minutes.
Unconstrained Face Detection and Open-Set Face Recognition Challenge
The dataset includes polarimetric, RGB and depth automotive (on the road) data.