TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

Symbrain

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksImages, MRI, Medical

agda2train

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

open-compass/CriticBench

[Dataset on HF] [Project Page] [Subjective LeaderBoard] [Objective LeaderBoard]

1 papers0 benchmarksTexts

prompt-opin-summ

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksTexts

opin-pref

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarksTexts

Unreal-RS

Rolling Shutter dataset captured with Unreal engine

1 papers0 benchmarks

QTL (Quantitative Trait Locus)

We present QTL, a real-life DS-NER application in the animal science domain. The entity type to recognize is “trait”, an important task in the construction of genotype-phenotype databases for advancing livestock genomics research and breeding methodologies. Different from previous DS-NER datasets, where entities consist of many proper nouns, trait entities consist of descriptive expressions. For example, in the sentence “There are several reports on candidate gene analysis of subcutaneous fat depth traits using similar Wagyu × Limousin F2populations”, “subcutaneous fat depth traits” is a trait entity that needs to be identified.

1 papers0 benchmarks

CityFFD 3D Urban Wind Simulation - Niigata (High-resolution 3D urban microclimate simulation in multiple wind directions)

This is the full dataset for the paper Fourier neural operator for real-time simulation of 3D dynamic urban microclimate. The dataset of 3D urban wind simulation data of Niigata is generated from CityFFD. A total of 1200 steps of wind simulation were executed. The dataset contains four wind directions of data. Data for the west and north winds include all 1200 simulation steps. Data for the east and south winds include the last 50 steps of the simulation. Each step of the data is a 200 * 200 * 150 array with 32-bit precision and is stored as a numpy file.

1 papers0 benchmarks3D

Educational Grade School Math (EGSM)

Educational Grade School Math (EGSM) contains 2,093 question/answer pairs generated by MATHWELL, a reference-free educational grade school math word problem generator that outputs a word problem and Program of Thought (PoT) solution based solely on an optional student interest, as introduced in MATHWELL: Generating Educational Math Word Problems Using Teacher Annotations. The question/answer pairs are verified by human experts. EGSM is the first teacher-annotated math word problem training dataset for LLMs.

1 papers0 benchmarksTexts

MATHWELL Human Annotation Dataset

The MATHWELL Human Annotation Dataset contains 5,084 synthetic word problems and answers generated by MATHWELL, a reference-free educational grade school math word problem generator released in MATHWELL: Generating Educational Math Word Problems Using Teacher Annotations, and comparison models (GPT-4, GPT-3.5, Llama-2, MAmmoTH, and LLEMMA) with expert human annotations for solvability, accuracy, appropriateness, and meets all criteria (MaC). Solvability means the problem is mathematically possible to solve, accuracy means the Program of Thought (PoT) solution arrives at the correct answer, appropriateness means that the mathematical topic is familiar to a grade school student and the question's context is appropriate for a young learner, and MaC denotes questions which are labeled as solvable, accurate, and appropriate. Null values for accuracy and appropriateness indicate a question labeled as unsolvable, which means it cannot have an accurate solution and is automatically inappropria

1 papers0 benchmarksTexts

CNFOOD-241-Chen

CNFOOD-241 Contains a dataset of 241 Chinese dishes with 191,811 images. There are 170843 images in the training set and 20943 images in the validation set. All images are resized to 600x600. As some of the images in the dataset are from ChineseFoodNet, they are not supported for commercial use. CNFOOD-241-Chen is the CNFOOD-241 dataset spilt with the list introduced in the paper "Res-VMamba: Fine-Grained Food Category Visual Classification Using Selective State Space Models with Deep Residual Learning," which has random split as train, val, test three parts.

1 papers1 benchmarksImages

CNFOOD-241

Contains a dataset of 241 Chinese dishes with 191,811 images. There are 170843 images in the training set and 20943 images in the validation set. All images are resized to 600x600. As some of the images in the dataset are from ChineseFoodNet, they are not supported for commercial use.

1 papers0 benchmarksImages

AE (answer correctness dataset)

Click to add a brief description of the dataset (Markdown and LaTeX enabled).

1 papers0 benchmarks

LEARNING STYLE IDENTIFICATION (Learning Style Identification Using Semi-Supervised Self-Taught Labeling)

The dataset was collected from two courses offered on the University of Jordan's E-learning Portal during the second semester of 2020, namely "Computer Skills for Humanities Students" (CSHS) and "Computer Skills for Medical Students" (CSMS). Over the sixteen-week duration of each course, students participated in various activities such as reading materials, video lectures, assignments, and quizzes. To preserve student privacy, the log activity of each student was anonymized. Data was aggregated from multiple sources, including the Moodle learning management system and the student information system, and consolidated into a single database. The dataset contains information on the number of learners and events for each course, as well as their launch and end dates. CSHS had 1749 learners and 1,139,810 events from January 21, 2020 to May 20, 2020, while CSMS had 564 learners and 484,410 events during the same period. The dataset is based on the Filder and Silverman learning style model (F

1 papers0 benchmarksActions, Texts, Tracking

Radio-Freqency Ultrasound volume dataset for pre-clinical liver tumors

A total of 227 cross sectional images (20 x 54 mm with a resolution of 289 x 648 pixels) of hind-leg xenograft tumors from 29 mice were obtained with 1mm step-wise movement of the array mounted on a manual positioning device. The whole tumor volume was acquired using a diagnostic ultrasound system with a 10 MHz linear transducer and 50 MHz sampling.

1 papers0 benchmarks3D, Biomedical, Images, Medical

BaitBuster-Bangla: A Comprehensive Dataset for Clickbait Detection in Bangla with Multi-Feature and Multi-Modal Analysis

The dataset contains a total of 253,070 records, with 18 features. The features are categorized into four different types: Metadata, Primary Data, Engagement Stats, and Label. Under the Metadata category contains basic information about the channel and video, such as their unique identifiers, date and time of publication, and thumbnail URLs. The Primary Data category contains information about the title and description of the video. The "Processed" columns refer to the cleaned data after denoising, deduplication and debiased for further analysis. The Engagement Stats category contains data on user engagement metrics for each video. The Label category contains predefined auto labels, human annotated labels, and AI generated pseudo labels. Auto labels are labels that are automatically derived based on a review of their titles, descriptions, and thumbnails over time. Channels with consistently misleading, exaggerated, or sensationalized content were labeled as clickbait. Those focusing on

1 papers0 benchmarksTabular, Texts

PolymerAbstracts

A dataset of 750 polymer abstracts annotated with the entity types: POLYMER, POLYMER_CLASS, PROPERTY_VALUE, PROPERTY_NAME, MONOMER, ORGANIC_MATERIAL, INORGANIC_MATERIAL, and MATERIAL_AMOUNT. This data set was annotated for training a named entity recognition model for these entity types.

1 papers0 benchmarks

Character-LLM Data

This is the training datasets for Character-LLM, which contains nine characters experience data used to train Character-LLMs.

1 papers0 benchmarks

Satellite

The Satellite dataset forms a practical VFL scenario for location identification based on satellite imagery. Each AOI, with its unique location identifier, is captured by 16 satellite visits. Assuming each visit is carried out by a distinct satellite organization, these organizations aim to collectively train a model to classify the land type of the location without sharing original images. The Satellite dataset encompasses four land types as labels, namely Amnesty POI (4.8%), ASMSpotter (8.9%), Landcover (61.3%), and UNHCR (25.0%), making the task a 4-class classification problem of 3927 locations, containing 62,832 images across 16 parties, simulating a practical VFL scenario of collaborative location identification via multiple satellites.

1 papers0 benchmarksImages

CAER-Dynamic

13,201 clips from 79 TV shows. Each video clip was manually annotated with six emotion categories, including “anger”, “disgust”, “fear”, “happy”, “sad”, and “surprise“, as well as “neutral”.

1 papers4 benchmarks
PreviousPage 489 of 1000Next