Generalizing Verifiable Instruction Following

Valentina Pyatkin, Saumya Malik, Victoria Graf, Hamish Ivison, Shengyi Huang, Pradeep Dasigi, Nathan Lambert, Hanna Hajishirzi

arXiv:2507.02833 · 2026-08-15 공개 · arXiv · PDF

reinforcement-learning language-models instruction-following rlvr generalization training-data constraint-verification verifiable-constraints

Abstract

A crucial factor for successful human and AI interaction is the ability of language models or chatbots to follow human instructions precisely. A common feature of instructions are output constraints like ``only answer with yes or no"or ``mention the word `abrakadabra'at least 3 times"that the user adds to craft a more useful answer. Even today's strongest models struggle with fulfilling such constraints. We find that most models strongly overfit on a small set of verifiable constraints from the benchmarks that test these abilities, a skill called precise instruction following, and are not able to generalize well to unseen output constraints. We introduce a new benchmark, IFBench, to evaluate precise instruction following generalization on 58 new, diverse, and challenging verifiable out-of-domain constraints. In addition, we perform an extensive analysis of how and on what data models can be trained to improve precise instruction following generalization. Specifically, we carefully design constraint verification modules and show that reinforcement learning with verifiable rewards (RLVR) significantly improves instruction following. In addition to IFBench, we release 29 additional new hand-annotated training constraints and verification functions, RLVR training prompts, and code.

한국어 요약

한 줄 요약

IFBench라는 새로운 벤치마크를 제시하며, RLVR 기반 학습이 정밀 지시사항 준수 성능을 15% 이상 향상시킨다.

핵심 기여도

핵심 아이디어

기존 모델은 IFEval과 같은 작은 벤치마크에 과적합되어 있고, 새로운 제약 조건에 대한 일반화 능력이 부족하다. 이에 따라, IFBench라는 새로운 벤치마크를 제안하며, 모델이 다양한 제약 조건을 일반화할 수 있도록 RLVR 기반 학습을 도입한다. RLVR는 제약 조건을 검증 가능한 보상으로 변환하여 강화 학습을 통해 모델을 훈련시킨다. 예를 들어, "특정 단어를 3번 이상 포함"하는 제약 조건은 Python 함수로 자동 검증 가능하며, 이를 보상으로 사용하여 학습을 유도한다. 이는 기존 지시사항 준수 훈련 방식과는 차별화된 접근법이다.

기술적 접근법

주요 결과

의의 및 한계

IFBench는 기존 모델이 제약 조건에 과적합된 문제를 드러내며, 정밀 지시사항 준수의 일반화 능력을 평가하는 새로운 기준을 제시한다. RLVR와 GRPO를 통한 훈련은 기존 지시사항 준수 성능을 유지하면서 새로운 제약 조건에 대한 일반화를 향상시킨다. 그러나 연구는 검증 가능한 제약 조건에만 집중하며, 실제 사용 환경에서의 비검증 가능한 제약 조건은 다루지 않는다. 또한 일부 제약 조건이 인위적일 수 있어, 자연스러운 제약 조건에 대한 연구가 필요하다.

실용적 활용

IFBench와 RLVR 기반 학습은 챗봇, 고객 지원 시스템, 자동화된 문서 생성 등 정밀 지시사항 준수가 필수적인 산업에 적용 가능하다. 특히, 사용자 지정 제약 조건을 자동으로 처리하는 AI 시스템 개발에 기여할 수 있다.