跳到主要导航 跳到搜索 跳到主要内容

E-EVAL: A Comprehensive Chinese K-12 Education Evaluation Benchmark for Large Language Models

  • Jinchang Hou
  • , Chang Ao
  • , Haihong Wu
  • , Xiangtao Kong
  • , Zhigang Zheng
  • , Daijia Tang
  • , Chengming Li
  • , Xiping Hu
  • , Ruifeng Xu
  • , Shiwen Ni*
  • , Min Yang*
  • *此作品的通讯作者
  • Shenzhen Institute of Advanced Technology
  • University of Science and Technology of China
  • Southern University of Science and Technology
  • Union Information
  • Shenzhen MSU-BIT University
  • Harbin Institute of Technology Shenzhen

科研成果: 书/报告/会议事项章节会议稿件同行评审

摘要

The rapid development of Large Language Models (LLMs) has led to their increasing utilization in Chinese K-12 education. Despite the growing integration of LLMs and education, the absence of a dedicated benchmark for evaluating LLMs within this domain presents a pressing concern. Consequently, there is an urgent need for a comprehensive natural language processing benchmark to precisely assess the capabilities of various LLMs in Chinese K-12 education. In response, we introduce E-EVAL, the first comprehensive evaluation benchmark specifically tailored for Chinese K-12 education. E-EVAL comprises 4,351 multiple-choice questions spanning primary, middle, and high school levels, covering a diverse array of subjects. Through meticulous evaluation, we find that Chinese-dominant models often outperform English-dominant ones, with many exceeding GPT 4.0. However, most struggle with complex subjects like mathematics. Additionally, our analysis indicates that most Chinese-dominant LLMs do not achieve higher scores at the primary school level compared to the middle school level, highlighting the nuanced relationship between proficiency in higher-order and lower-order knowledge domains. Furthermore, experimental results highlight the effectiveness of the Chain of Thought (CoT) technique in scientific subjects and Few-shot prompting in liberal arts. Through E-EVAL, we aim to conduct a rigorous analysis delineating the strengths and limitations of LLMs in educational applications, thereby contributing significantly to the advancement of Chinese K-12 education and LLMs.

源语言英语
主期刊名The 62nd Annual Meeting of the Association for Computational Linguistics
主期刊副标题Findings of the Association for Computational Linguistics, ACL 2024
编辑Lun-Wei Ku, Andre Martins, Vivek Srikumar
出版商Association for Computational Linguistics (ACL)
7753-7774
页数22
ISBN(电子版)9798891760998
DOI
出版状态已出版 - 2024
已对外发布
活动Findings of the 62nd Annual Meeting of the Association for Computational Linguistics, ACL 2024 - Hybrid, Bangkok, 泰国
期限: 11 8月 202416 8月 2024

出版系列

姓名Proceedings of the Annual Meeting of the Association for Computational Linguistics
ISSN(印刷版)0736-587X

会议

会议Findings of the 62nd Annual Meeting of the Association for Computational Linguistics, ACL 2024
国家/地区泰国
Hybrid, Bangkok
时期11/08/2416/08/24

指纹

探究 'E-EVAL: A Comprehensive Chinese K-12 Education Evaluation Benchmark for Large Language Models' 的科研主题。它们共同构成独一无二的指纹。

引用此