19,997 machine learning datasets
19,997 dataset results
Hawk Annotation Dataset includes language descriptions specifically for anomaly scenes in seven existing video anomaly datasets. These seven datasets include a variety of anomalous scenarios such as crime (UCF-Cirme), campus (ShanghaiTech and CUHK Avenue), pedestrian walkways (UCSD Ped1 and Ped2), traffic (DoTA), and human behavior (UBnormal). With the support of these visual scenarios, this dataset can perform comprehensive fine-tuning for various abnormal scenarios, being closer to open-world scenarios.
A large-scale, egocentric, multimodal, and context-aware dataset of human demonstrations of social navigation.
This dataset contains 65 DFIs acquired from patients with POAG at the University of Iowa Hospitals and Clinics. DFIs were acquired using a 30° Zeiss fundus camera (Niemeijer et al 2011). The images were centered on the optic disc. The original DFIs resolution was 2392 × 2048. In order to benchmark LUNet on this dataset, the black border of the DFIs were padded to a squared resolution of 2048 × 2048 pixels and then resized to a 1444 × 1444 pixels resolution. From the resulting DFIs, 15 optic disc-centered DFIs were randomly selected to form the second external test set. No other additional metadata were provided in the open source dataset.
The forbidden question dataset they build (based on two previous works) contains 160 questions from 160 violated categories. In addition, they also provide the corresponding targets – which is useful for some jailbreak methods, such as GCG.
HALvest is a textual dataset comprising 17 billion tokens in 56 languages and 13 domains.
HALvest-Geometric is a subset of HALvest: an academic citation network with 238,397 disambiguated authors and 18,662,037 scholarly papers.
Bluesky Social Dataset Pollution of online social spaces caused by rampaging d/misinformation is a growing societal concern. However, recent decisions to reduce access to social media APIs are causing a shortage of publicly available, recent, social media data, thus hindering the advancement of computational social science as a whole. To address this pressing issue, we present a large, high-coverage dataset of social interactions and user-generated content from Bluesky Social.
Heritage Pointcloud Instance Collection dataset, acquired from two large buildings and annotated at a point-wise semantic level based on existent BIM models. Devid Campagnolo, Elena Camuffo, Umberto Michieli, Paolo Borin, Simone Milani and Andrea Giordano, "Fully Automated Scan-to-BIM via Point Cloud Instance Segmentation", In Proceedings of the International Conference on Image Processing (ICIP) 2023.
The Synthesis Blockchain Intrusion Detection System dataset
Enhancing Financial Market Predictions: Causality-Driven Feature Selection This paper introduces FinSen dataset that revolutionizes financial market analysis by integrating economic and financial news articles from 197 countries with stock market data. The dataset’s extensive coverage spans 15 years from 2007 to 2023 with temporal information, offering a rich, global perspective 160,000 records on financial market news. Our study leverages causally validated sentiment scores and LSTM models to enhance market forecast accuracy and reliability.
Here is the forbidden question dataset (based on two previous works), it contains 160 questions from 160 violated categories. In addition, authors also provide the corresponding target -- which is useful for some jailbreak methods, such as GCG.
Here is the forbidden question dataset (based on two previous works), it contains 160 questions from 160 violated categories. In addition, authors also provide the corresponding target -- which is useful for some jailbreak methods, such as GCG.
Expository-Prose-V1 is a collection of specially-curated corpora gathered from diverse sources, ranging from research papers (arXiv) to European Parliament proceedings (EuroParl). It has been specially filtered and curated for the quality of text, depth of reasoning and breadth of knowledge to faciliate an effective pre-train. It was used to pre-train 1.5-Pints, a small but powerful Large Language Model developed by the Pints Research Team.
Noise of Web (NoW) is a challenging noisy correspondence learning (NCL) benchmark for robust image-text matching/retrieval models. It contains 100K image-text pairs consisting of website pages and multilingual website meta-descriptions (98,000 pairs for training, 1,000 for validation, and 1,000 for testing). NoW has two main characteristics: without human annotations and the noisy pairs are naturally captured. The source image data of NoW is obtained by taking screenshots when accessing web pages on mobile user interface (MUI) with 720 $\times$ 1280 resolution, and we parse the meta-description field in the HTML source code as the captions. In NCR (predecessor of NCL), each image in all datasets were preprocessed using Faster-RCNN detector provided by Bottom-up Attention Model to generate 36 region proposals, and each proposal was encoded as a 2048-dimensional feature. Thus, following NCR, we release our the features instead of raw images for fair comparison. However, we can not just
Table-MovieLens1M (TML1M) is a relational table dataset derived from the classical MovieLens1M dataset. It consists of three tables: users, movies, and ratings. Notably, the movie table has been enriched with more comprehensive features. Additionally, the dataset defines a standard classification task focused on predicting user age ranges.
Table-LastFm2K (TLF2K) is a relational table dataset derived from the classical LastFM2K dataset. It contains three tables: artists, user_artists, and user_friends. Notably, the artists table has been enhanced with more detailed features, and the tags for each artist have been streamlined. The dataset also provides a standard classification task for music genre classification of artists.
Table-ACM12K (TACM12K) is a relational table dataset derived from the ACM heterogeneous graph dataset. It includes four tables: papers, authors, citations, and writings. The paper table features attributes such as year, title, and abstract, while the author table includes name and affiliation details. Additionally, some feature completion has been performed for the papers. The dataset also defines a standard classification task for predicting the conference to which a paper belongs.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
The dataset used in this study is obtained from the U.S. Department of Transportation, Bureau of Transportation Statistics from January 2019– August 2023 . It contains 32 attributes related to planned f light date-time, airline, planned origin and destination, cancellation and diversion status, overall delay, and delay due to individual components (carrier, weather, NAS, security, late aircraft), among others.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).