TY - GEN
T1 - Automatic Evaluating Scientific Reviews Through Meta-reviewer’s Lens
T2 - 14th National CCF Conference on Natural Language Processing and Chinese Computing, NLPCC 2025
AU - Liu, Shu Hang
AU - Lan, Tian
AU - Zhang, Yun He
AU - Liu, Ran
AU - Huang, Heyan
AU - Xu, Chen
AU - Wu, Zhijing
AU - Mao, Xian Ling
N1 - Publisher Copyright:
© The Author(s), under exclusive license to Springer Nature Singapore Pte Ltd. 2026.
PY - 2026
Y1 - 2026
N2 - Peer review, essential for scientific manuscript assessment, faces efficiency challenges due to increasing submissions, demanding reliable automatic metrics for review quality evaluation. The currently most widely used word-overlap and embedding-based metrics using single peer review references show significant discrepancies with human judgments. To address these issues, we propose a novel automatic metric for evaluating the quality of peer reviews, designed to analyze the quality of atomic review opinions (AROs) using large language models, named ReviewScore. Besides, we construct a high-quality benchmark through human annotations to reliably measure ReviewScore and peer review generation models, called ReviewEval. The reference reviews in ReviewEval are crafted by integrating valuable opinions from multiple reviewers based on meta review, making them more holistic and objective. Experimental results show ReviewScore’s superior alignment with human judgments compared to existing metrics. Besides, using ReviewEval, we comprehensively re-evaluate peer review generation models and conduct detailed analysis, revealing several key insights. The ReviewEval benchmark and the toolkit for ReviewScore will be publicly released.
AB - Peer review, essential for scientific manuscript assessment, faces efficiency challenges due to increasing submissions, demanding reliable automatic metrics for review quality evaluation. The currently most widely used word-overlap and embedding-based metrics using single peer review references show significant discrepancies with human judgments. To address these issues, we propose a novel automatic metric for evaluating the quality of peer reviews, designed to analyze the quality of atomic review opinions (AROs) using large language models, named ReviewScore. Besides, we construct a high-quality benchmark through human annotations to reliably measure ReviewScore and peer review generation models, called ReviewEval. The reference reviews in ReviewEval are crafted by integrating valuable opinions from multiple reviewers based on meta review, making them more holistic and objective. Experimental results show ReviewScore’s superior alignment with human judgments compared to existing metrics. Besides, using ReviewEval, we comprehensively re-evaluate peer review generation models and conduct detailed analysis, revealing several key insights. The ReviewEval benchmark and the toolkit for ReviewScore will be publicly released.
KW - Automatic Eva luation
KW - Large Language Model
KW - Peer Review
UR - https://www.scopus.com/pages/publications/105046013535
U2 - 10.1007/978-981-95-3346-6_38
DO - 10.1007/978-981-95-3346-6_38
M3 - Conference contribution
AN - SCOPUS:105046013535
SN - 9789819533459
T3 - Lecture Notes in Computer Science
SP - 496
EP - 507
BT - Natural Language Processing and Chinese Computing - 14th National CCF Conference, NLPCC 2025, Proceedings
A2 - Mao, Xian-Ling
A2 - Ren, Zhaochun
A2 - Yang, Muyun
PB - Springer Science and Business Media Deutschland GmbH
Y2 - 7 August 2025 through 9 August 2025
ER -