Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review

Haowen Li, Yoichi Ishibashi, M. Oyamada

arXiv:2607.22553 · 2026-07-28 공개 · arXiv · PDF

llm peer-review llm-generated reviewer-guidelines automated-review conference-guidelines rubric-scoring human-judgments

Abstract

Peer review is an essential process in scientific research, yet the growing workload has made its automation increasingly necessary. In this study, we analyze how different types of reviewer guidelines, such as official conference guidelines and reviewer-imitating ones generated from high-quality human reviews using LLMs, affect automated peer review. Our experiments show that official conference guidelines produce review results most consistent with human judgments, suggesting that evaluation criteria refined through conference practice serve as effective guidance for automated reviewing as well. In contrast, reviewer-imitating guidelines were generally less effective than official conference guidelines. Furthermore, enforcing strict rubric-style scoring consistently degraded performance, highlighting the importance of allowing subjective and holistic scoring.

한국어 요약

한 줄 요약

LLM 기반 자동 학술 심사에서 심사 가이드라인 설계가 결과에 미치는 영향을 평가한 연구.

핵심 기여도

핵심 아이디어

이 연구는 학술 심사 자동화의 질을 높이기 위해, LLM이 사용하는 심사 가이드라인의 설계 방식이 결과에 미치는 영향을 분석한다. 공식 학회 가이드라인은 학회 운영 경험을 바탕으로 정제된 평가 기준을 제공하므로, 인간 심사와 가장 유사한 결과를 도출할 수 있다. 반면, LLM이 생성한 심사자 흉내 가이드라인은 인간의 심사 방식을 정확히 반영하지 못해 효과가 떨어진다. 또한, 점수 부여 시 엄격한 체크리스트 방식은 창의성과 종합적 판단을 무시하게 되어 성능을 저하시킨다.

기술적 접근법

주요 결과

의의 및 한계

이 연구는 LLM 기반 자동 학술 심사 시, 가이드라인 설계가 결과에 미치는 영향을 정량적으로 평가한 점에서 학술적 의의가 있다. 특히, 인간 심사와의 일관성을 높이기 위해 공식 가이드라인 사용을 권장하며, 점수 부여 방식의 유연성도 강조한다. 한계로는, 실험은 특정 데이터셋과 평가 척도에 기반했기 때문에 일반화 가능성에 제약이 있을 수 있다.

실용적 활용

이 연구는 학술 심사 자동화 시스템 개발자들에게 가이드라인 설계 방향을 제시하며, 특히 학회 운영자나 편집위원회가 LLM 도구를 활용할 때 유용한 참고 자료가 될 수 있다. 또한, 점수 부여 방식의 유연성은 학술 심사 외에도 교육 평가, 채용 면접 등 다양한 분야에서 적용 가능하다.