Skip to main navigation Skip to search Skip to main content

Automatic Evaluating Scientific Reviews Through Meta-reviewer’s Lens: A Reliable Benchmark for Peer Review Generation

  • Beijing Institute of Technology

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Peer review, essential for scientific manuscript assessment, faces efficiency challenges due to increasing submissions, demanding reliable automatic metrics for review quality evaluation. The currently most widely used word-overlap and embedding-based metrics using single peer review references show significant discrepancies with human judgments. To address these issues, we propose a novel automatic metric for evaluating the quality of peer reviews, designed to analyze the quality of atomic review opinions (AROs) using large language models, named ReviewScore. Besides, we construct a high-quality benchmark through human annotations to reliably measure ReviewScore and peer review generation models, called ReviewEval. The reference reviews in ReviewEval are crafted by integrating valuable opinions from multiple reviewers based on meta review, making them more holistic and objective. Experimental results show ReviewScore’s superior alignment with human judgments compared to existing metrics. Besides, using ReviewEval, we comprehensively re-evaluate peer review generation models and conduct detailed analysis, revealing several key insights. The ReviewEval benchmark and the toolkit for ReviewScore will be publicly released.

Original languageEnglish
Title of host publicationNatural Language Processing and Chinese Computing - 14th National CCF Conference, NLPCC 2025, Proceedings
EditorsXian-Ling Mao, Zhaochun Ren, Muyun Yang
PublisherSpringer Science and Business Media Deutschland GmbH
Pages496-507
Number of pages12
ISBN (Print)9789819533459
DOIs
Publication statusPublished - 2026
Externally publishedYes
Event14th National CCF Conference on Natural Language Processing and Chinese Computing, NLPCC 2025 - Urumqi, China
Duration: 7 Aug 20259 Aug 2025

Publication series

NameLecture Notes in Computer Science
Volume16103 LNAI
ISSN (Print)0302-9743
ISSN (Electronic)1611-3349

Conference

Conference14th National CCF Conference on Natural Language Processing and Chinese Computing, NLPCC 2025
Country/TerritoryChina
CityUrumqi
Period7/08/259/08/25

Keywords

  • Automatic Eva luation
  • Large Language Model
  • Peer Review

Fingerprint

Dive into the research topics of 'Automatic Evaluating Scientific Reviews Through Meta-reviewer’s Lens: A Reliable Benchmark for Peer Review Generation'. Together they form a unique fingerprint.

Cite this