large-language-models reward-hacking peer-review evaluation-protocols llm-rewriting ai-reviewers rhetorical-sensitivity scientific-evaluation
Abstract
As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions. We construct a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions. Two LLM rewriters transform six rhetorical dimensions in opposing directions, and five LLM reviewers evaluate the resulting manuscripts under standard and strict protocols. We also test joint, recursive, and reviewer-guided rewriting. Our results show that rhetorical sensitivity is structured rather than uniform. Evidence framing and novelty stance produce the largest positive-negative contrasts in overall assessment, with scope framing forming a weaker second tier; the remaining dimensions have smaller or less stable effects. This hierarchy persists across human-assessed quality levels, but score movement depends strongly on the AI reviewer's original score: lower scores tend to rise, higher scores tend to fall, and directional contrasts are clearest in the middle ranges. More elaborate workflows do not reliably yield larger gains. Joint rewriting is strongly rewriter-dependent, reviewer guidance does not consistently outperform an unguided second pass, and repeated rewriting yields diminishing, configuration-dependent returns. Across conditions, the rewriter primarily determines the separation between opposing variants, whereas the reviewer determines the magnitude and sign of their score effects. Strict review lowers mean OA by 1.36 points without consistently changing rhetorical sensitivity. These findings identify when rhetorical presentation influences AI scientific review and motivate evaluation systems robust to content-preserving variation in scientific writing.
한국어 요약
한 줄 요약
AI 리뷰어의 판단이 과학적 내용을 유지하면서도 수사적 표현에 따라 달라질 수 있음을 밝힌 연구.
핵심 기여도
- 4,200개의 ICLR 2026 제출 논문을 기반으로 6개의 수사적 차원을 양극성으로 재작성한 실험 설계.
- GPT-5.5, Opus 4.8 등 2개의 LLM 리라이터와 Gemini 3.5 FL, Qwen 3.5 F 등 5개의 LLM 리뷰어를 사용한 비교 실험.
- "evidence framing"과 "novelty stance"가 전체 평가에 가장 큰 영향을 미치며, "scope framing"이 그 다음으로 영향력 있음.
- "strict review"는 평균 OA 점수를 1.36점 낮추지만 수사적 민감도에는 일관된 변화 없음.
핵심 아이디어
AI 리뷰어는 과학적 내용이 동일하더라도 수사적 표현 방식에 따라 평가 점수가 달라질 수 있다. 이 연구는 6개의 수사적 차원 — 주장 및 혁신성 태도, 범위 및 일반화, 수치적 증거 구성, 기여 구조, 기술적 표현 수준, 어휘 및 문법 복잡도 — 를 양극성으로 재작성하여 AI 리뷰어의 민감도를 측정했다. 특히, "evidence framing"과 "novelty stance"가 가장 큰 영향을 미치며, 리라이터와 리뷰어의 모델 종류, 평가 프로토콜에 따라 민감도가 달라진다는 점이 핵심 통찰이다. 이는 AI 리뷰 시스템이 단순히 평균 점수나 일관성만으로 평가될 수 없다는 것을 시사한다.
기술적 접근법
- **데이터셋**: ICLR 2026 제출 논문 120개를 기반으로 4,200개의 원고 생성.
- **리라이터 모델**: GPT-5.5 (Codex CLI), Opus 4.8 (Claude Code) 사용.
- **리뷰어 모델**: Gemini 3.5 FL, Qwen 3.5 F, GPT-5 mini, GPT-5.5, Sonnet 5.
- **수사적 차원**: 6개의 차원을 양극성으로 재작성.
- **작업 유형**: 단일 차원 재작성(2,880개), 공동 재작성(240개), 재귀적 공동 재작성(480개), 리뷰어 지도형 재작성(480개).
- **평가 프로토콜**: "Standard", "Strict" 두 가지.
- **총 리뷰 기록**: 42,396개.
- **비용**: $29,165.89 (API 호출 기준).
주요 결과
- "evidence framing"과 "novelty stance"가 전체 평가에 가장 큰 영향을 미침.
- "scope framing"은 그 다음으로 영향력 있음.
- AI 리뷰어의 초기 점수에 따라 점수 변화가 달라짐: 낮은 점수는 상승, 높은 점수는 하락, 중간 점수에서 가장 명확한 변화.
- "strict review"는 평균 OA 점수를 1.36점 낮추지만 수사적 민감도는 일관되지 않음.
- 공동 재작성은 리라이터에 강하게 의존하며, 반복 재작성은 감소하는 수익을 보임.
의의 및 한계
이 연구는 AI 리뷰 시스템이 수사적 표현에 민감하게 반응할 수 있음을 실증적으로 보여주며, 과학적 내용과 표현 방식을 명확히 분리하는 것이 중요함을 강조한다. 특히, 리라이터와 리뷰어 모델, 평가 프로토콜에 따라 결과가 달라지므로, AI 리뷰 시스템은 단일 평가 기준이 아닌 다중 모델과 조건에서의 "수사적 강건성"을 평가해야 한다는 점이 학술적 의의이다. 한계로는 실험 대상이 ICLR 2026 제출 논문에 한정되었으며, 실제 학술 리뷰 시스템과의 직접적인 비교는 명시되지 않았다.
실용적 활용
이 연구는 과학 논문 작성 및 리뷰 과정에서 AI 도구를 사용할 때 수사적 표현이 평가에 미치는 영향을 고려해야 함을 시사한다. 특히, AI 리뷰어를 도입하는 학술 저널이나 연구 기관은 수사적 민감도를 고려한 평가 프로토콜을 설계해야 하며, 리라이터와 리뷰어 모델의 선택에 주의를 기울여야 한다.