19,997 machine learning datasets
19,997 dataset results
We assembled a benchmark of electronic component pinouts, PINS100, containing 100 common parts frequently used in circuits found on high-traffic electronic tutorial websites such as the ARDUINO PROJECT HUB and AUTODESK TINKERCAD CIRCUITS. Components range from 2 pins to 40 pins, and span a large assortment of part categories including passives (e.g. resistors/capacitors), input (e.g. switches), output (e.g. LEDs, motors, relays), sensors, integrated circuits, power regulators, logic (e.g. 7400- SERIES AND and OR gates), and microcontrollers (e.g ARDUINO, RASPBERRY PI ).
To assess a model’s ability to create microcontroller-driven electronic devices, we developed a benchmark, MICRO25, that includes 25 tasks intended for the common ARDUINO microcontroller ecosystem.. These tasks, shown in Table 2, span 5 core categories including: input, interface protocols, output, sensors, and logic. Each task is either tailored to test a specific fundamental competency required to build basic microcontroller-driven electronic devices, or the integration of several competencies into larger design flows.
Click to add a brief description of the dataset (Markdown and LaTeX enabled). D Provide: D* a high-level explanation of the dataset characteristics * explain motivations and summary of its content * potential use cases of the dataset --DASDSSDSDS
The goal of InfoLossQA is to generate a series of QA pairs that reveal to lay readers what information a simplified text lacks compared to its original.
TwBNT is the bot detection benchmark with automatic troll annotations.
🔍 About data: Our database is designed to be a key tool in the advancement of practical research and application in rotating machinery monitoring. It covers a variety of operating conditions and fault types, making the data applicable to a wide range of industrial scenarios.
a corpus for topic modeling in Albanian
Named Entity Recognition in Albanian
Sentiment Analysis of Movie Reviews in Albanian
The Flipkart Products Review Dataset is a collection of reviews provided by customers who purchased products from Flipkart, one of India’s leading e-commerce platforms. This dataset is valuable for sentiment analysis, which helps determine the emotional tone of customer reviews. Sentiment analysis is commonly used in e-commerce and other platforms to understand customer opinions and experiences with products.
This is a multilabel dataset used for Noise Identification purpose in the paper "A Comparative Analysis of Noise Reduction Methods in Sentiment Analysis on Noisy Bangla Texts" accepted in 2024 The 9th Workshop on Noisy and User-generated Text (W-NUT) collocated with EACL 2024.
TARA is a dataset for tool-augmented reward modeling, which includes comprehensive comparison data of human preferences and detailed tool invocation processes.
ConCon: Continually Confounded Dataset is a confounded visual dataset for continual learning.
HierText is the first dataset featuring hierarchical annotations of text in natural scenes and documents. The dataset contains 11639 images selected from the Open Images dataset, providing high quality word (~1.2M), line, and paragraph level annotations. Text lines are defined as connected sequences of words that are aligned in spatial proximity and are logically connected. Text lines that belong to the same semantic topic and are geometrically coherent form paragraphs. Images in HierText are rich in text, with average of more than 100 words per image.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
This datasets consists of challenging reasoning questions in multiple choice format.
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Click to add a brief description of the dataset (Markdown and LaTeX enabled).
Concept-1K contains 1023 novel concepts from six domains, including economy, culture, science and technology, environment, education, and health and medical. It has 16653 training-test QA pairs corresponding to 16653 knowledge points from 1023 concepts. It is proposed for evaluating the forgetting in large language models and the effectiveness of incremental learning algorithms.