TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

PINS100

We assembled a benchmark of electronic component pinouts, PINS100, containing 100 common parts frequently used in circuits found on high-traffic electronic tutorial websites such as the ARDUINO PROJECT HUB and AUTODESK TINKERCAD CIRCUITS. Components range from 2 pins to 40 pins, and span a large assortment of part categories including passives (e.g. resistors/capacitors), input (e.g. switches), output (e.g. LEDs, motors, relays), sensors, integrated circuits, power regulators, logic (e.g. 7400- SERIES AND and OR gates), and microcontrollers (e.g ARDUINO, RASPBERRY PI ).

1 papers0 benchmarksTexts

MICRO25

To assess a model’s ability to create microcontroller-driven electronic devices, we developed a benchmark, MICRO25, that includes 25 tasks intended for the common ARDUINO microcontroller ecosystem.. These tasks, shown in Table 2, span 5 core categories including: input, interface protocols, output, sensors, and logic. Each task is either tailored to test a specific fundamental competency required to build basic microcontroller-driven electronic devices, or the integration of several competencies into larger design flows.

1 papers0 benchmarksTexts

D-Nikud dataset

Click to add a brief description of the dataset (Markdown and LaTeX enabled). D Provide: D* a high-level explanation of the dataset characteristics * explain motivations and summary of its content * potential use cases of the dataset --DASDSSDSDS

1 papers0 benchmarks

InfoLossQA

The goal of InfoLossQA is to generate a series of QA pairs that reveal to lay readers what information a simplified text lacks compared to its original.

1 papers0 benchmarksTexts

TwBNT

TwBNT is the bot detection benchmark with automatic troll annotations.

1 papers0 benchmarks

DATA-PHM

🔍 About data: Our database is designed to be a key tool in the advancement of practical research and application in rotating machinery monitoring. It covers a variety of operating conditions and fault types, making the data applicable to a wide range of industrial scenarios.

1 papers0 benchmarks

AlbNews

a corpus for topic modeling in Albanian

1 papers0 benchmarks

AlbNER

Named Entity Recognition in Albanian

1 papers0 benchmarks

AlbMoRe

Sentiment Analysis of Movie Reviews in Albanian

1 papers0 benchmarks

Flipkart Products Review

The Flipkart Products Review Dataset is a collection of reviews provided by customers who purchased products from Flipkart, one of India’s leading e-commerce platforms. This dataset is valuable for sentiment analysis, which helps determine the emotional tone of customer reviews. Sentiment analysis is commonly used in e-commerce and other platforms to understand customer opinions and experiences with products.

1 papers0 benchmarks

NC-SentNoB (Noise Classification on SentNoB)

This is a multilabel dataset used for Noise Identification purpose in the paper "A Comparative Analysis of Noise Reduction Methods in Sentiment Analysis on Noisy Bangla Texts" accepted in 2024 The 9th Workshop on Noisy and User-generated Text (W-NUT) collocated with EACL 2024.

1 papers0 benchmarksTexts

TARA

TARA is a dataset for tool-augmented reward modeling, which includes comprehensive comparison data of human preferences and detailed tool invocation processes.

1 papers0 benchmarksTexts

ConCon Dataset (Continually Confounded Dataset)

ConCon: Continually Confounded Dataset is a confounded visual dataset for continual learning. 

1 papers0 benchmarks

HierText

HierText is the first dataset featuring hierarchical annotations of text in natural scenes and documents. The dataset contains 11639 images selected from the Open Images dataset, providing high quality word (~1.2M), line, and paragraph level annotations. Text lines are defined as connected sequences of words that are aligned in spatial proximity and are logically connected. Text lines that belong to the same semantic topic and are geometrically coherent form paragraphs. Images in HierText are rich in text, with average of more than 100 words per image.

1 papers5 benchmarks

DeepFigures Open (Extracting Scientific Figures with Distantly Supervised Neural Networks)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

IMU_blur

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

bigscience/P3 (bigscience/P3, split='ai2_arc_ARC_Challenge_pick_the_most_correct_option')

This datasets consists of challenging reasoning questions in multiple choice format.

1 papers0 benchmarksTexts

Pothole Dataset (Pothole_detection_inference.)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksImages

BdSLW60 (Bangla Word Level Sign Language Dataset)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

Concept-1K

Concept-1K contains 1023 novel concepts from six domains, including economy, culture, science and technology, environment, education, and health and medical. It has 16653 training-test QA pairs corresponding to 16653 knowledge points from 1023 concepts. It is proposed for evaluating the forgetting in large language models and the effectiveness of incremental learning algorithms.

1 papers0 benchmarksTexts
PreviousPage 487 of 1000Next