JaQuAD: Japanese Question Answering Dataset for Machine Reading Comprehension

ByungHoon So, Kyuhong Byun, Kyungwon Kang, Seongjin Cho

2022-02-03Reading Comprehension Question Answering Machine Reading Comprehension

Abstract

Question Answering (QA) is a task in which a machine understands a given document and a question to find an answer. Despite impressive progress in the NLP area, QA is still a challenging problem, especially for non-English languages due to the lack of annotated datasets. In this paper, we present the Japanese Question Answering Dataset, JaQuAD, which is annotated by humans. JaQuAD consists of 39,696 extractive question-answer pairs on Japanese Wikipedia articles. We finetuned a baseline model which achieves 78.92% for F1 score and 63.38% for EM on test set. The dataset and our experiments are available at https://github.com/SkelterLabsInc/JaQuAD.

Results

Task	Dataset	Metric	Value	Model
Question Answering	JaQuAD	Exact Match	63.38	BERT-Japanese
Question Answering	JaQuAD	F1	78.92	BERT-Japanese

Related Papers

From Roots to Rewards: Dynamic Tree Reasoning with RL2025-07-17 Enter the Mind Palace: Reasoning and Planning for Long-term Active Embodied Question Answering2025-07-17 Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter It2025-07-17 City-VLM: Towards Multidomain Perception Scene Understanding via Multimodal Incomplete Learning2025-07-17 Describe Anything Model for Visual Question Answering on Text-rich Images2025-07-16 Is This Just Fantasy? Language Model Representations Reflect Human Judgments of Event Plausibility2025-07-16 Warehouse Spatial Question Answering with LLM Agent2025-07-14 Evaluating Attribute Confusion in Fashion Text-to-Image Generation2025-07-09