VNHSGE: VietNamese High School Graduation Examination Dataset for Large Language Models

Dao Xuan-Quy, Le Ngoc-Bich, Vo The-Duy, Phan Xuan-Dung, Ngo Bac-Bien, Nguyen Van-Tien, Nguyen Thi-My-Thanh, Nguyen Hong-Phuoc

2023-05-20Reading Comprehension Question Answering Text Generation Question Rewriting Multiple-choice Visual Question Answering

Paper PDF Code(official)

Abstract

The VNHSGE (VietNamese High School Graduation Examination) dataset, developed exclusively for evaluating large language models (LLMs), is introduced in this article. The dataset, which covers nine subjects, was generated from the Vietnamese National High School Graduation Examination and comparable tests. 300 literary essays have been included, and there are over 19,000 multiple-choice questions on a range of topics. The dataset assesses LLMs in multitasking situations such as question answering, text generation, reading comprehension, visual question answering, and more by including both textual data and accompanying images. Using ChatGPT and BingChat, we evaluated LLMs on the VNHSGE dataset and contrasted their performance with that of Vietnamese students to see how well they performed. The results show that ChatGPT and BingChat both perform at a human level in a number of areas, including literature, English, history, geography, and civics education. They still have space to grow, though, especially in the areas of mathematics, physics, chemistry, and biology. The VNHSGE dataset seeks to provide an adequate benchmark for assessing the abilities of LLMs with its wide-ranging coverage and variety of activities. We intend to promote future developments in the creation of LLMs by making this dataset available to the scientific community, especially in resolving LLMs' limits in disciplines involving mathematics and the natural sciences.

Results

Task	Dataset	Metric	Value	Model
Question Answering	VNHSGE-English	Accuracy	92.4	Bing Chat
Question Answering	VNHSGE-English	Accuracy	79.2	ChatGPT
Question Answering	VNHSGE-History	Accuracy	88.5	Bing Chat
Question Answering	VNHSGE-History	Accuracy	56.5	ChatGPT
Question Answering	VNHSGE-Biology	Accuracy	69	Bing Chat
Question Answering	VNHSGE-Biology	Accuracy	58	ChatGPT
Question Answering	VNHSGE Mathematics	Accuracy	60	Bing Chat
Question Answering	VNHSGE Mathematics	Accuracy	58.8	ChatGPT
Question Answering	VNHSGE-Civic	Accuracy	85.5	Bing Chat
Question Answering	VNHSGE-Civic	Accuracy	70.5	ChatGPT
Question Answering	VNHSGE-Literature	Accuracy	68	ChatGPT
Question Answering	VNHSGE-Literature	Accuracy	56.8	Bing Chat
Question Answering	VNHSGE-Physics	Accuracy	66	Bing Chat
Question Answering	VNHSGE-Physics	Accuracy	61	ChatGPT
Question Answering	VNHSGE-Geography	Accuracy	85.5	Bing Chat
Question Answering	VNHSGE-Geography	Accuracy	61.5	ChatGPT
Question Answering	VNHSGE-Chemistry	Accuracy	52.5	Bing Chat
Question Answering	VNHSGE-Chemistry	Accuracy	48	ChatGPT

VNHSGE: VietNamese High School Graduation Examination Dataset for Large Language Models

Abstract

Results

Related Papers

VNHSGE: VietNamese High School Graduation Examination Dataset for Large Language Models

Abstract

Results

Related Papers