MultiChallenge: A Realistic Multi-Turn Conversation Evaluation Benchmark Challenging to Frontier LLMs

Ved Sirdeshmukh, Kaustubh Deshpande, Johannes Mols, Lifeng Jin, E. Cardona, Dean Lee, Jeremy Kritz, Willow E. Primack, Summer Yue, Chen Xing

arXiv:2501.17399 · 2026-07-27 공개 · arXiv · PDF

llm-evaluation instruction-following llm-as-judge multi-turn-conversation frontier-llms context-allocation in-context-reasoning human-llm-interaction

Abstract

We present MultiChallenge, a pioneering benchmark evaluating large language models (LLMs) on conducting multi-turn conversations with human users, a crucial yet underexamined capability for their applications. MultiChallenge identifies four categories of challenges in multi-turn conversations that are not only common and realistic among current human-LLM interactions, but are also challenging to all current frontier LLMs. All 4 challenges require accurate instruction-following, context allocation, and in-context reasoning at the same time. We also develop LLM as judge with instance-level rubrics to facilitate an automatic evaluation method with fair agreement with experienced human raters. Despite achieving near-perfect scores on existing multi-turn evaluation benchmarks, all frontier models have less than 50% accuracy on MultiChallenge, with the top-performing Claude 3.5 Sonnet (June 2024) achieving just a 41.4% average accuracy.

한국어 요약

한 줄 요약

MultiChallenge는 최첨단 LLM들이 다중 턴 대화에서 실패하는 4가지 실제적 도전 과제를 평가하는 새로운 벤치마크이다.

핵심 기여도

핵심 아이디어

MultiChallenge는 기존 다중 턴 평가가 실제 인간-LLM 대화의 복잡성을 충분히 반영하지 못한다는 문제를 해결하기 위해 설계되었다. 기존 평가 벤치마크는 명시적 지시 따르기 위주로 구성되어 실제 대화에서 요구되는 다양한 능력을 평가하지 못한다. 본 연구는 instruction retention, inference memory, reliable versioned editing, self-coherence라는 4가지 도전 과제를 정의하여, LLM이 대화 맥락을 유지하면서 정확한 추론과 지시 준수를 동시에 수행하는 능력을 평가한다. 특히, self-coherence는 LLM이 인간의 의견에 무조건적으로 동의하는 'sycoophancy'를 피하는 능력을 평가하는 핵심 지표이다.

기술적 접근법

주요 결과

의의 및 한계

MultiChallenge는 기존 다중 턴 평가가 실제 대화의 복잡성을 반영하지 못하는 문제를 해결하며, LLM의 실제 대화 능력을 정확히 평가할 수 있는 새로운 기준을 제시한다. 특히, LLM-as-judge와 인스턴스 루브릭을 결합한 자동 평가 시스템은 인간 평가와 높은 일치율을 보이며, 효율적인 평가 도구로 활용 가능하다. 그러나, MultiChallenge은 6개의 최첨단 모델이 실패한 예제만 포함하기 때문에, 이 모델들에 대해 상대적으로 편향된 결과가 나올 수 있다는 한계가 있다.

실용적 활용

MultiChallenge는 대화형 AI의 실제 대화 능력을 평가하는 데 유용하며, 특히 고객 서비스, 교육, 가상 보조 등 인간-LLM 대화가 필요한 산업에서 모델 선택과 개선에 활용될 수 있다. 또한, 인스턴스 레벨 자동 평가 시스템은 대규모 LLM 평가를 효율적으로 수행할 수 있는 도구로 활용 가능하다.