19,997 machine learning datasets
19,997 dataset results
TITANIC-FGS is the first domain knowledge-enhanced, instruction-following dataset specifically designed for the Remote Sensing Fine-Grained Ship Classification (RS-FGSC) task. It simulates human-like step-by-step decision-making to train vision-language models (VLMs) for interpretable and accurate ship classification.
Large Shape and Texture dataset (LAS&T) is a giant dataset of shapes and textures for tasks of visual shapes and textures identification and retrieval from single image.
49 videos of gastrojejunostomy procedure
This dataset was created as part of the Master's thesis titled "Multi-Class Depression Detection Through Tweets Using Artificial Intelligence." It contains tweets labeled for five types of depression (Bipolar, Major, Psychotic, Atypical, and Postpartum) using lexicons verified by psychiatrists. Click to add a brief description of the dataset (Markdown and LaTeX enabled).
This package contains the data and the reported results for the manuscript: Keo B, Li B, Younis W (2025) Measuring trade costs and analyzing the determinants of trade growth between Cambodia and major trading partners: 1993–2019. PLoS ONE 20(1): e0311754. https://doi.org/10.1371/journal.pone.0311754
A comprehensive object-instance ReID dataset with multiple indoor object instances under varying lighting conditions.
A real world dataset for benchmarking global localization in complex indoor environments.
A ProcTHOR created synthetic dataset for benchmarking global localization in complex indoor environments.
Government fiscal policies, particularly annual union budgets, exert significant influence on financial markets. However, real-time analysis of budgetary impacts on sector-specific equity performance remains methodologically challenging and largely unexplored. This study proposes a framework to systematically identify and rank sectors poised to benefit from India's Union Budget announcements. The framework addresses two core tasks: (1) multi-label classification of excerpts from budget transcripts into 81 predefined economic sectors, and (2) performance ranking of these sectors. Leveraging a comprehensive corpus of Indian Union Budget transcripts from 1947 to 2025, we introduce BASIR (Budget-Assisted Sectoral Impact Ranking), an annotated dataset mapping excerpts from budgetary transcripts to sectoral impacts.
In the realm of social media, understanding and predicting post reach is a significant challenge. Our paper presents a Crowd Reaction AssessMent (CReAM) task designed to estimate if a given social media post will receive more reaction than another, a particularly essential task for digital marketers and content writers. We introduce the Crowd Reaction Estimation Dataset (CRED), consisting of pairs of tweets from The White House with comparative measures of retweet count.
Predicting stock market prices following corporate earnings calls remains a significant challenge for investors and researchers alike, requiring innovative approaches that can process diverse information sources. This study investigates the impact of corporate earnings calls on stock prices by introducing a multi-modal predictive model. We leverage textual data from earnings call transcripts, along with images and tables from accompanying presentations, to forecast stock price movements on the trading day immediately following these calls. To facilitate this research, we developed the MiMIC (Multi-Modal Indian Earnings Calls) dataset, encompassing companies representing the Nifty 50, Nifty MidCap 50, and Nifty Small 50 indices. The dataset includes earnings call transcripts, presentations, fundamentals, technical indicators, and subsequent stock prices. We present a multimodal analytical framework that integrates quantitative variables with predictive signals derived from textual and v
We present two multi-modal datasets, one for Main Board IPOs, and the other for Small and Medium Enterprises (SME) IPOs. It consists of various features relating to the company going for IPOs, and other macroeconomic factors. The objective is to estimate the direction and under pricing with respect to opening, high and closing prices of stocks on the IPOlisting day.
https://github.com/dialogue-evaluation/RuOpinionNE-2024
Open-source dataset
Collected by cleaning data from daily Xinwen Lianbo transcripts over the past three months and processing it using reverse engineering techniques.
A synthetic dataset from an automobile manufacturer datasource.
The FairTranslate Dataset includes 2,418 sentence pairs, each centered around an occupation, designed to assess gender expression and translation in English-French contexts. Each English sentence appears in three gender variants (male, female, inclusive), allowing for direct counterfactual comparisons. This structure supports fairness evaluations and helps analyze how models handle grammatical gender, inclusive forms, and coreference resolution in translation.
The Balanced Choice of Plausible Alternatives dataset is a benchmark for training machine learning models that are robust to superficial cues/spurious correlations. The dataset extends the COPA dataset(Roemmele et al. 2011) with mirrored instances that mitigate against token-level superficial cues in the original COPA answers. The superficial cues in the original COPA datasets result from an unbalanced token distribution between the correct and the incorrect answer choices, i.e., some tokens appear more in the correct choices than the incorrect ones. Balanced COPA equalizes the token distribution by adding mirrored instances with identical answer choices but different labels. The details about the creation of Balanced COPA and the implementation of the baselines are available in the paper.
Two versions of the dataset are offered: one is the full dataset used to train the models in our paper, and the other is a mini dataset for easier examination. Both versions include raw and postprocessed subsets of peeling, wiping and lifting. The raw videos of the tactile dataset used for generate the PCA embedding are also provided.
FBIS-22M is the largest field boundary instance segmentation dataset to date, featuring over 22 million labeled field instances across more than 672 000 high-resolution satellite image patches. It includes imagery from 0.25m to 10m resolution, sourced from multiple satellites and covering diverse geographic regions, enabling robust training for scalable agricultural vision models.