TasksSotADatasetsPapersMethodsSubmitAbout
Papers With Code 2

A community resource for machine learning research: papers, code, benchmarks, and state-of-the-art results.

Explore

Notable BenchmarksAll SotADatasetsPapersMethods

Community

Submit ResultsAbout

Data sourced from the PWC Archive (CC-BY-SA 4.0). Built by the community, for the community.

Datasets

19,997 machine learning datasets

Filter by Modality

  • Images3,275
  • Texts3,148
  • Videos1,019
  • Audio486
  • Medical395
  • 3D383
  • Time series298
  • Graphs285
  • Tabular271
  • Speech199
  • RGB-D192
  • Environment148
  • Point cloud135
  • Biomedical123
  • LiDAR95
  • RGB Video87
  • Tracking78
  • Biology71
  • Actions68
  • 3d meshes65
  • Tables52
  • Music48
  • EEG45
  • Hyperspectral images45
  • Stereo44
  • MRI39
  • Physics32
  • Interactive29
  • Dialog25
  • Midi22
  • 6D17
  • Replay data11
  • Financial10
  • Ranking10
  • Cad9
  • fMRI7
  • Parallel6
  • Lyrics2
  • PSG2

19,997 dataset results

File S5 (tree file of Form IV rubiscos (RLPs) at 70% dereplication)

This tree was generated as indicated above in the methods. The model chosen by the algorithm was LG+F+R8.

1 papers0 benchmarks

File S6 (tree file of Form I rubiscos at 85% dereplication)

This tree was generated as indicated above in the text for file S3 using model LG+F+G.

1 papers0 benchmarks

File S7 (tree file of Form I rubiscos at 85% dereplication LG+R8)

This tree was generated as indicated above in the methods. The model chosen by the algorithm was LG+R8.

1 papers0 benchmarks

ValNov Subtask B

Validity and Novelty are determined in a comparative setting between two conclusions at a time. For Validity and Novelty possible labels are "Conclusion 1 is better", "tie" and "Conclusion 2 is better", for Validity and Novelty respectively.

1 papers12 benchmarksTexts

InstructPix2Pix Image Editing Dataset

A dataset for image editing containing >450k samples of:

1 papers0 benchmarksImages, Texts

RGB Arabic Alphabet Sign Language (AASL) dataset

RGB Arabic Alphabet Sign Language (AASL) dataset

1 papers1 benchmarksImages

Bochrum movement data

The linked repository holds data from a controlled single-obstacle avoidance experiment recorded in a motion laboratory. If you are using this data, please cite the following sources for the data

1 papers0 benchmarks

HOI-SDC (Setting for Double Challenge of Human-Object Interaction Detection)

In order to avoid the training process of the model being influenced by a portion of HOI classes with a very small number of instances, we remove some of the HOI classes containing a very small number of instances and HOI classes with no interaction from the training \textbf{S}et for the \textbf{D}ouble \textbf{C}hallenge. Finally, there are total 321 HOI classes, 74 object classes and 93 action classes. The training and testing set contain 37,155 and 9,666 images, respectively.

1 papers0 benchmarks

HC3

The HC3 (Human ChatGPT Comparison Corpus) dataset consists of nearly 40K questions and their corresponding human/ChatGPT answers. The motivation for this dataset was to study ChatGPT's answers in contrast to human's answers. The questions range from a wide variety of domains, including open-domain, financial, medical, legal, and psychological areas.

1 papers0 benchmarks

STEDUCOV: A DATASET ON STANCE DETECTION IN TWEETS TOWARDS ONLINE EDUCATION DURING COVID-19 PANDEMIC

StEduCov, a dataset annotated for stances toward online education during the COVID-19 pandemic. StEduCov has 17,097 tweets gathered over 15 months, from March 2020 to May 2021, using Twitter API. The tweets are manually annotated into agree, disagree or neutral classes. We used a set of relevant hashtags and keywords. Specifically, we utilised a combination of hashtags, such as '#COVID 19' or '#Coronavirus' with keywords, such as 'education', 'online learning', 'distance learning' and 'remote learning'. To ensure high annotation quality, three different annotators annotated each tweet and at least one of the reviewers from three judges revised it. They were guided by some instructions, such as that in the case of disagree class, there should be a clear negative statement about online education or its impact. Also, if the tweet is negative but refers to other people (e.g. 'my children hate online learning').

1 papers2 benchmarks

WMT-SLT

We provide separate training, development and test data. The training data is available right away. The development and test data will be released in several stages, starting with a release of the development sources only.

1 papers0 benchmarksRGB Video, Texts

Govdocs1

GovDocs is a corpus of nearly 1 million documents that are freely available for research and may be, to the best of the authors' knowledge, freely redistributed. These documents were obtained by performing searches for words randomly chosen from the Unix dictionary, numbers randomly chosen between 1 and 1 million, and randomized combinations of the two, for documents of specified file types that resided on web servers in the .gov domain using the Yahoo an Google search engines. The documents are representative of a diverse sample of real-world files of various formats produced by a variety of tools, including any malware that may be present in the files. Therefore, the corpus has been used in digital forensics, malware analysis, computer vision, and natural language processing research.

1 papers0 benchmarks

Complete data from the Barro Colorado 50-ha plot: 423617 trees, 35 years

The 50-ha plot at Barro Colorado Island was initially demarcated and fully censused in 1982, and has been fully censused 7 times since, every 5 years from 1985 through 2015. Every measurement of every stem over 8 censuses is included in this archive. Most users will need only the 8 R Analytical Tables in the format tree, which come here zipped together into a single archive (bci.tree.zip), plus the single R Species Table.

1 papers0 benchmarksEnvironment

FZ queries (FindZebra queries)

A set of 248 search queries annotated with the correct diagnosis. The diagnosis is referenced with a Concept Unique Identifier (CUI). In a retrieval setting, the task consists of retrieving an article from the FindZebra corpus with a CUI that matches the query CUI.

1 papers0 benchmarks

MTTN

MTTN is a large scale derived and synthesized dataset built with on real prompts and indexed with popular image-text datasets like MS-COCO, Flickr, etc. MTTN consists of over 2.4M sentences that are divided over 5 stages creating a combination amounting to over 12M pairs, along with a vocab size of consisting more than 300 thousands unique words that creates an abundance of variations.

1 papers0 benchmarksTexts

UICaption

UICaption is a dataset of 114k UI images paired with descriptions of their functionality. It is designed for the tasks of UI action entailment, instruction-based UI image retrieval, grounding referring expressions, and UI entity recognition.

1 papers0 benchmarksImages, Texts

PushWorld

PushWorld is an environment with simplistic physics that requires manipulation planning with both movable obstacles and tools. It contains more than 200 PushWorld puzzles in PDDL and in an OpenAI Gym environment.

1 papers0 benchmarksEnvironment

REN-20k Dataset

Reader Emotion News 20k Dataset

1 papers0 benchmarksTexts

EHT (The English Headline Treebank)

The English Headline Treebank (EHT) is an English headline treebank of 1,055 manually annotated and adjudicated universal dependency (UD) syntactic dependency trees to encourage research in improving NLP pipelines for English headlines.

1 papers0 benchmarksTexts

ConsInv Dataset

ConsInv is a stereo RGB + IMU dataset designed for Dynamic SLAM testing and contains two subsets:

1 papers0 benchmarksImages, RGB Video, Stereo
PreviousPage 453 of 1000Next