InternReviewer & InternAdvocate: Objective Reward and Evaluation for Agentic Reinforcement Learning in Peer Review and Rebuttal

Xuerui Su, Liya Guo, Qizhi Pei, Qipeng Guo, Zhongbo Tian, Lijun Wu, Kai Chen, Zun Wang

arXiv:2608.28612 · 2026-09-02 공개 · arXiv · PDF

agentic-rl semantic-alignment peer-review reference-anchored verification-mechanism scholarly-dataset rebuttal-generation scholarly-agents

Abstract

Generating professional scholarly content, such as peer reviews and rebuttals, requires an intricate synergy between domain reasoning and factual grounding. This work presents a comprehensive framework for the development and evaluation of specialized scholarly agents, InternReviewer and InternAdvocate. We first establish a large-scale, high-quality scholarly dataset and integrate a high-efficiency arXiv retrieval tool to enable active evidence gathering. To optimize these agents, we implement an agentic Reinforcement Learning (RL) paradigm driven by a unified objective metric and reward system. This system avoids the biases of subjective model-based judging by employing multi-dimensional criteria, including reference-anchored semantic alignment, structural compliance, and a strict verification mechanism that cross-checks citations against real-time interaction logs to eliminate hallucinations. Experimental results demonstrate that agents trained within this closed-loop framework exhibit significant improvements in reasoning depth and citation accuracy.

한국어 요약

한 줄 요약

InternReviewer와 InternAdvocate는 강화학습 기반의 학술 리뷰 및 반박 생성 에이전트로, 객관적 보상 설계와 실시간 검증을 통해 신뢰성과 정확도를 향상시킨다.

핵심 기여도

핵심 아이디어

기존 학술 리뷰 생성 시스템은 주관적 평가 기준에 의존하거나, 사실 기반 검증이 부족한 문제가 있었다. 본 연구는 이 문제를 해결하기 위해 **Objective Reward Design**을 도입한다. 이는 **reference-anchored semantic alignment**, **structural compliance**, **strict verification mechanism**을 기반으로 구성되며, 특히 인용문을 실시간 인터랙션 로그와 비교하여 hallucination을 제거한다.
또한, **Agentic Reinforcement Learning**을 통해 학습 과정에서 외부 도구(arXiv 검색)와 상호작용하며, 정확하고 신뢰성 있는 리뷰를 유도한다. 이는 기존 SFT 방식과 달리, 단순 스타일 모방이 아닌 사실 기반 논리 구성을 학습하도록 유도한다.

기술적 접근법

주요 결과

의의 및 한계

본 연구는 학술 리뷰 생성 자동화 분야에서 **객관적 평가 기준**을 제시하며, **LLM-as-a-judge의 주관성 문제**를 해결하는 중요한 기여를 한다. 특히, **실시간 검증 기반의 인용 검증**은 학술 리뷰의 신뢰성을 높이는 데 기여한다.
하지만, **긴 문맥 처리**는 여전히 한계로 남아 있으며, 수학적 내용이나 다량의 그림이 포함된 논문은 모델의 컨텍스트 윈도우와 파서 성능을 초과할 수 있다. 또한, **도구 호출 시스템**은 arXiv 외의 소스(예: Google Scholar)로 확장되어야 보다 포괄적인 검색이 가능하다.

실용적 활용

InternReviewer와 InternAdvocate는 **학술 컨퍼런스의 리뷰 과부하 문제**를 완화하고, **일관되고 신뢰성 있는 피드백 제공**에 활용될 수 있다. 또한, **AI Scientist 생태계 내의 내부 평가 구성 요소**로 사용되어, 가설 생성, 실험, 글쓰기, 평가 사이의 피드백 루프를 닫는 데 기여할 수 있다.