WebWalker: Benchmarking LLMs in Web Traversal

Jialong Wu, Wenbiao Yin, Yong Jiang, Zhenglin Wang, Zekun Xi, Runnan Fang, Deyu Zhou, Pengjun Xie, Fei Huang

arXiv:2501.07572 · 2026-07-27 공개 · arXiv · PDF

retrieval-augmented multi-agent llm-benchmark information-extraction llm-framework web-traversal webwalkerqa explore-critic

Abstract

Retrieval-augmented generation (RAG) demonstrates remarkable performance across tasks in open-domain question-answering. However, traditional search engines may retrieve shallow content, limiting the ability of LLMs to handle complex, multi-layered information. To address it, we introduce WebWalkerQA, a benchmark designed to assess the ability of LLMs to perform web traversal. It evaluates the capacity of LLMs to traverse a website's subpages to extract high-quality data systematically. We propose WebWalker, which is a multi-agent framework that mimics human-like web navigation through an explore-critic paradigm. Extensive experimental results show that WebWalkerQA is challenging and demonstrates the effectiveness of RAG combined with WebWalker, through the horizontal and vertical integration in real-world scenarios.

한국어 요약

한 줄 요약

WebWalkerQA는 LLM이 웹을 체계적으로 탐색하는 능력을 평가하는 다단계 벤치마크이며, WebWalker라는 탐색-비판 다중 에이전트 프레임워크를 제안한다.

핵심 기여도

핵심 아이디어

기존 검색 엔진은 단일 층의 정보만 제공해 LLM이 복잡한 웹 구조를 탐색하는 능력을 제한한다. 이를 해결하기 위해 WebWalkerQA는 LLM이 주어진 웹사이트 내 하위 페이지를 체계적으로 탐색해 정보를 추출하는 능력을 평가하는 벤치마크를 제안한다. WebWalker는 인간의 웹 탐색을 모방한 다중 에이전트 프레임워크로, 탐색 에이전트(ReAct 기반)와 비판 에이전트로 구성되어 메모리 관리와 추론을 수행한다. 이는 수직적 탐색을 통해 정보를 깊이 있게 추출하는 새로운 패러다임을 제시한다.

기술적 접근법

주요 결과

의의 및 한계

WebWalkerQA는 LLM이 실제 웹 환경에서 정보를 체계적으로 추출하는 능력을 평가하는 첫 번째 다단계 벤치마크로, RAG와의 통합 가능성을 제시한다. 특히, 수직적 탐색을 통해 정보의 깊이를 확보하는 방식은 정보 검색 시스템의 확장성과 신뢰도를 높이는 데 기여할 수 있다. 그러나, 일부 복잡한 탐색 시나리오에서는 여전히 계획 및 추론 능력이 부족하며, 긴 텍스트 생성의 정확도 평가가 어려운 한계가 있다.

실용적 활용

WebWalker는 웹 기반 정보 추출, 고객 지원 시스템, 온라인 교육 플랫폼 등에서 LLM의 탐색 능력을 향상시키는 데 활용될 수 있다. 특히, 정보가 다층적으로 구성된 정부, 학술, 기업 웹사이트에서의 활용성이 높다.