InFoBench: Evaluating Instruction Following Ability in Large Language Models

Yiwei Qin, Kaiqiang Song, Yebowen Hu, Wenlin Yao, Sangwoo Cho, Xiaoyang Wang, Xuansheng Wu, Fei Liu, Pengfei Liu, Dong Yu

arXiv:2401.03601 · 2026-07-27 공개 · arXiv · PDF

llm-evaluation instruction-following llm-benchmark llm-compliance drfr-metric infobench annotation-sources constraint-categories

Abstract

This paper introduces the Decomposed Requirements Following Ratio (DRFR), a new metric for evaluating Large Language Models' (LLMs) ability to follow instructions. Addressing a gap in current methodologies, DRFR breaks down complex instructions into simpler criteria, facilitating a detailed analysis of LLMs' compliance with various aspects of tasks. Alongside this metric, we present InFoBench, a benchmark comprising 500 diverse instructions and 2,250 decomposed questions across multiple constraint categories. Our experiments compare DRFR with traditional scoring methods and explore annotation sources, including human experts, crowd-sourced workers, and GPT-4. The findings demonstrate DRFR's higher reliability and the effectiveness of using GPT-4 as a cost-efficient annotator. The evaluation of several advanced LLMs using this framework reveals their strengths and areas needing improvement, particularly in complex instruction-following. This study contributes a novel metric and benchmark, offering insights for future LLM development and evaluation.

한국어 요약

한 줄 요약

InFoBench는 LLM의 지시사항 따르기 능력을 평가하기 위한 새로운 메트릭 DRFR과 500개의 지시사항을 포함한 벤치마크를 제시한다.

핵심 기여도

핵심 아이디어

기존 평가 방법은 지시사항을 통합적으로 평가하는 데 한계가 있어, InFoBench는 DRFR라는 새로운 메트릭을 통해 각 지시사항을 세부 기준으로 분해함으로써 모델의 특정 영역별 성능을 정확히 분석할 수 있도록 설계되었다. 이는 기존의 전체 점수 기반 평가(DS)보다 해석성이 높고, 특히 복잡한 지시사항(Hard Set)에서 신뢰도가 더 높다. 또한, GPT-4를 활용한 자동 평가 방식은 전문가와 유사한 정확도를 유지하면서도 시간과 비용을 크게 절감할 수 있다.

기술적 접근법

주요 결과

의의 및 한계

InFoBench와 DRFR는 LLM의 지시사항 처리 능력을 세부적으로 평가할 수 있는 체계적인 프레임워크를 제공하며, 특히 복잡한 지시사항에서의 모델 성능을 분석하는 데 유용하다. 또한, GPT-4를 활용한 자동 평가 방식은 대규모 평가를 효율적으로 수행할 수 있는 가능성을 제시한다. 그러나, DRFR는 지시사항 분해 과정에서 인간의 주관이 개입될 수 있으며, 일부 모델은 수치적 이해나 언어적 해석에서 여전히 개선이 필요하다는 한계가 있다.

실용적 활용

InFoBench와 DRFR는 LLM 개발자들이 모델의 지시사항 처리 능력을 정확히 평가하고, 복잡한 사용 시나리오에서의 성능을 개선하는 데 활용될 수 있다. 특히, GPT-4를 활용한 자동 평가 방식은 대규모 모델 평가를 저비용으로 수행할 수 있어, 산업 및 연구 분야에서 실용적 가치가 높다.