19,997 machine learning datasets
19,997 dataset results
[Real or Fake] : Fake Job Description Prediction This dataset contains 18K job descriptions out of which about 800 are fake. The data consists of both textual information and meta-information about the jobs. The dataset can be used to create classification models which can learn the job descriptions which are fraudulent.
The GMDCSA dataset contains 16 ADL (not fall) activities and 16 Fall activities. The GMDCSA dataset has been created by performing the fall and the ADL activities by a single subject wearing a different set of clothes. The web camera of a laptop (HP 348 G5 Laptop: Core i5 8th Gen/8 GB/512 GB SSD/Windows 10) was used to capture the activities.
Collected more than 10,854 samples (4,354 malware and 6,500 benign) from several sources.
The SMS Spam Collection is a public set of SMS labeled messages that have been collected for mobile phone spam research.
This ExAIS_SMS Spam dataset was a project conducted at the Federal University of Agriculture, Abeokuta, Nigeria with the aim of building an indigenous SMS Spam corpus with African-English context.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Synthetic Question Answering dataset in Serbian, acquired by automatic translation of SQuAD.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
The $\textbf{360+x}$ dataset is a large-scale database that emphasizes a comprehensive multifaceted understanding of daily scenes. It provides diverse viewpoints and data modalities to emulate how humans obtain daily information in real-world scenarios. It includes 232 scene examples, each with an average duration of 6.2 minutes, spanning across 28 scene categories (comprising 15 indoor scenes and 13 outdoor scenes).
NeurIT Dataset is open-sourced for public research usage. It is collected using the customized robotic platform across three buildings. We collect the training, validation, and test-seen sets in Building A, and build the test-seen and test-unseen set in Building B and C. During data collection, the robot moves at varying speeds up to the maximum value (1.5m/s). The dataset contains 110 sequences, totaling around 15 hours of tracking data that corresponds to a travel distance of about 33.7 km. Each sequence of data lasts 6~10 minutes, containing both IMU data (acceleration, gyroscope, magnetometer) and the ground truth trajectory. The ratio of the training set, validation set, test-seen set, and test-unseen set is 15:3:3:4.
The dataset concerns ADL activities performed in a smart home environment in an Interwoven manner.
The Innodata Red Teaming Prompts aims to rigorously assess models’ factuality and safety. This dataset, due to its manual creation and breadth of coverage, facilitates a comprehensive examination of LLM performance across diverse scenarios.
The Innodata Red Teaming Prompts aims to rigorously assess models’ factuality and safety. This dataset, due to its manual creation and breadth of coverage, facilitates a comprehensive examination of LLM performance across diverse scenarios.
The Innodata Red Teaming Prompts aims to rigorously assess models’ factuality and safety. This dataset, due to its manual creation and breadth of coverage, facilitates a comprehensive examination of LLM performance across diverse scenarios.
LLM-generated output for compiling PDDL-domains and problems (5 scenarios, 5 runs per scenario): Specs 5 scenarios 5 trials.json - original json file Domain.pddl and Problem.pddl for each run are parsed out of the original json file. Plan.json - plan (if succeeded generating it)
This data set comprises 22 fundus images with their corresponding manual annotations for the blood vessels, separated as arteries and veins. It also include labels for glaucomatous / healthy, differentiating between normal tension glaucoma (NAG) and primary open angle glaucoma (POAG).
The dataset concerns toy tasks that a human should teach to a robot. The number of task repetitions is limited in the dataset since the human should demonstrate the task to the robot only a few times.
This dataset comprises video files (converted into tif format) that depict glomerular activation in mice. The activation was recorded as the response for 35 monomolecular odors. Wide-field 1-photon calcium imaging was recorded at a framerate of 100 Hz, in Thy1-GCaMP6f mice implanted with cranial windows over the olfactory bulb. Mice were head-fixed during imaging, with monomolecular odors presented in a randomized sequence for 2 seconds apiece during each trial.
The AASL-Clear dataset is a collection of RGB images featuring Arabic alphabet sign Language gestures with backgrounds removed. Each image in this dataset showcases clear, isolated hand gestures, allowing for precise recognition and analysis of Arabic sign language alphabets. With transparent backgrounds, this dataset provides a clean and focused resource for training deep learning models in the domain of Arabic sign language recognition and classification.
Overview The LaMini Dataset is an instruction dataset generated using h2ogpt-gm-oasst1-en-2048-falcon-40b-v2. It is designed for instruction-tuning pre-trained models to specialize them in a variety of downstream tasks.