19,997 machine learning datasets
19,997 dataset results
PhysNLU is a collection of 4 core datasets related to sentence classification, ordering, and coherence of physics explanations based on related tasks. Each dataset comprises explanations extracted from Wikipedia including derivations and mathematical language.
ArtImage is a synthetic dataset of articulated object models of 5 categories from PartNet-Mobility for articulated object tasks in category level.
MetaEval is a collection of 101 NLP tasks. It consists of 101 tasks in a benchmark that can be used for future probing and transfer learning.
This is a dataset for Bengali Captioning from Images.
PaSa is a dataset to train Machine Learning algorithms to automate the highlighting of patent paragraphs with semantic annotations. It consists of 150k samples obtained by traversing USPTO patents over a decade
BanglaEmotion is a manually annotated Bangla Emotion corpus, which incorporates the diversity of fine-grained emotion expressions in social-media text. More fine-grained emotion labels are considered such as Sadness, Happiness, Disgust, Surprise, Fear and Anger - which are, according to Paul Ekman (1999), the six basic emotion categories. For this task, a large amount of raw text data are collected from the user’s comments on two different Facebook groups (Ekattor TV and Airport Magistrates) and from the public post of a popular blogger and activist Dr. Imran H Sarker. These comments are mostly reactions to ongoing socio-political issues and towards the economic success and failure of Bangladesh. A total of 32923 comments are scraped from the three sources aforementioned above. Out of these, a total of 6314 comments were annotated into the six categories. The distribution of the annotated corpus is as follows:
Manuals and test set.
The Python dataset introduced in the Parallel Corpus paper (A Parallel Corpus of Python Functions and Documentation Strings for Automated Code Documentation and Code Generation), commonly used for evaluating automated code summarization.
The Java dataset introduced in Hybrid-DeepCom (Deep code comment generation with hybrid lexical and syntactical information), commonly used to evaluate automated code summarization. It is basically a further version of DeepCom-Java.
The UFPR-ADMR-v2 dataset contains 5,000 dial meter images obtained on-site by employees of the Energy Company of Paraná (Copel), which serves more than 4M consuming units in the Brazilian state of Paraná. The images were acquired with many different cameras and are available in the JPG format with 320×640 or 640×320 pixels (depending on the camera orientation). More details are available in our paper.
A small benchmark dataset for e-scooter rider detection task, and a trained model to support the detection of e-scooter riders from RGB images collected from natural road scenes.
In this repository you can find all the elaborate results that were used for the simulated evaluation of an innovative, optimized for real-life use, STC-based, multi-robot Coverage Path Planning (mCPP) algorithm. For this evaluation were introduced in "Apostolidis, S. D., Kapoutsis, P. C., Kapoutsis, A. C., & Kosmatopoulos, E. B. (2022). Cooperative multi-UAV coverage mission planning platform for remote sensing applications. Autonomous Robots, 1-28." 20 ROIs, of different shapes and areas, that may include obstacles inside them. These ROIs along with some benchmark results can be found here: https://github.com/savvas-ap/cpp-simulated-evaluations
PerPaDa is a Persian paraphrase dataset that is collected from users' input in a plagiarism detection system.
Multi-Language Vocabulary Evaluation Data Set (MuLVE) is a dataset consisting of vocabulary cards and real-life user answers, labeled indicating whether the user answer is correct or incorrect.
FIG-Loneliness (FIne-Grained Loneliness) is a dataset collected by using Reddit posts in two young adult-focused forums and two loneliness related forums consisting of a diverse age group. Annotations by trained human annotators for binary and fine-grained loneliness classifications of the posts are provided.
HS-BAN is a binary class hate speech (HS) dataset in Bangla language consisting of more than 50,000 labeled comments, including 40.17% hate and rest are non hate speech.
Synthetic visual inspection data of structural elements in bridges. The data is generated using the OpenIPDM toolbox "Generate Synthetic Dataset". For further details about the data generation and the properties of the dataset, refer to the software manual at https://github.com/CivML-PolyMtl/OpenIPDM/blob/main/Help
Symmetry-OOD is a dataset for symmetry perception by deep neural networks.
About the study This study was exploring the landscape of interpersonal conflicts during code review in following areas: - how these conflicts look like - what role do they play in software development - what are their consequences - what factors do play role in their appearance and severity - what strategies can be used to prevent and manage conflicts
Extensible Event Stream (XES) software event log obtained through instrumenting the NASA CEV class using the tool available at {https://svn.win.tue.nl/repos/prom/XPort/}. This event log contains method-call level events describing a single run of an exhaustive unit test suite for the Crew Exploration Vehicle (CEV) example available and documented at {http://babelfish.arc.nasa.gov/trac/jpf/wiki/projects/jpf-statechart} (trac) {http://babelfish.arc.nasa.gov/hg/jpf/jpf-statechart} (mercurial repository). Note that the life-cycle information in this log corresponds to method call (start) and return (complete), and captures a method-call hierarchy. We attached a slightly preprocessed variant of this event log, where the execution of each unit test method is represented as a separate trace.