FinanceComplexQA: Benchmarking Agentic Reasoning on Industrial-grade Financial Documents

Xianfu Cheng, Shiwei Zhang, Jiyu Zhao, Jian Yang, Xinyuan Wang, Ming Zhou, Weixiao Zhou, Xiangyuan Guan, Xiang Li, Zhenhe Wu, Ziyi Ni, Zhoujun Li, Bingjing Xu

arXiv:2607.19238 · 2026-07-24 공개 · arXiv · PDF

question-answering agentic-reasoning multi-hop-reasoning rag-systems financial-documents document-synthesis financial-analysis agent-as-a-judge

Abstract

Agentic Reasoning has become a transformative force in financial analysis due to its ability to integrate large-scale information and generate reliable and accurate content. However, when handling complex real-world problems, different agents still show significant performance variation. In this work, we design Finance-LaTeX SKILL, a skill for synthesizing financial documents with complex layouts based on expert knowledge. Using an agent workflow built on this skill, we generate 2,000 professional financial documents along with 6,000 high-quality question-answer pairs. To evaluate the overall capability of agents, we introduce FinanceComplexQA, a comprehensive open-ended generation benchmark for financial documents that closely resembles real-world scenarios. It contains 2,026 deep research tasks targeting 1009 financial documents. FinanceComplexQA has 8 key features: bilingual support; coverage of six mainstream scenarios and seven tasks; expert-level document reasoning questions; deep research of complex layouts; relatively stable and permanent reference answers; and precise evaluation through an Agent-as-a-Judge with multiple evaluation metrics. Using FinanceComplexQA, we conduct a comprehensive evaluation of leading RAG systems and agentic reasoning tools for financial document QA. Through identifying and analyzing failure cases, we provide an in-depth study of their capabilities in numerical computation, multi-hop reasoning, content summarization, and industry analysis.

한국어 요약

한 줄 요약

FinanceComplexQA는 1009개의 실제 금융 문서를 기반으로 2,026개의 깊은 연구 질문을 포함한 금융 문서 QA 벤치마크로, Agent-as-a-Judge 평가를 통해 정확도와 추론력을 평가한다.

핵심 기여도

핵심 아이디어

FinanceComplexQA는 금융 분석에서 필요한 다중 단계 추론과 복잡한 문서 구조를 반영한 QA 벤치마크로, 기존 QA 데이터셋이 단일 문서나 템플릿 기반 질문에 제한된 문제를 해결한다. 핵심 아이디어는 **Dual-Context Reasoning**과 **Cross-Layout Evidence Aggregation**이다. Dual-Context는 문서 내 정보와 금융 도메인 지식을 결합하여 추론을 강화하고, Cross-Layout는 텍스트, 표, 차트 등을 통합 분석하도록 요구한다. Finance-LaTeX SKILL은 전문가 지식과 자동 문서 생성을 결합하여 2,000개의 문서와 6,000개의 QA 쌍을 생성하며, 이는 벤치마크 확장의 기반을 제공한다.

기술적 접근법

주요 결과

의의 및 한계

FinanceComplexQA는 금융 분석에서 필요한 복합적 추론과 문서 구조 이해를 평가하는 첫 번째 대규모 벤치마크로, 금융 AI 에이전트의 신뢰성 향상에 기여한다. 특히, Agent-as-a-Judge 평가와 Cross-Layout 추론은 실제 금융 분석과 유사한 환경을 구축한다. 그러나 현재 시스템은 수치 계산, 다중 단계 추론, 문서 구조 이해에서 여전히 15~20%의 성능 격차를 보이며, 이는 금융 AI의 실용화를 위한 주요 과제로 남는다. 또한, 일부 시나리오에서는 정답률이 50% 미만으로, 데이터셋의 다변성과 시스템의 일반화 능력 개선이 필요하다.

실용적 활용

FinanceComplexQA는 금융 기관의 투자 분석, 리스크 평가, 보고서 생성 등에 활용 가능한 AI 시스템 개발에 기여할 수 있다. 특히, 금융 AI 에이전트의 수치 계산 정확도와 다중 문서 추론 능력을 향상시켜 투자 의사결정의 신뢰성을 높이는 데 유용하다. 또한, 금융 교육 및 연구 분야에서 실제 금융 문서를 기반으로 한 QA 훈련 데이터로도 활용 가능하다.