FinBen: A Holistic Financial Benchmark for Large Language Models
Qianqian Xie, Weiguang Han, Zhengyu Chen, Ruoyu Xiang, Xiao Zhang, Yueru He, Mengxi Xiao, Dong Li, Yongfu Dai, Duanyu Feng, Yijing Xu, Haoqiang Kang, Zi-Zhou Kuang, Chenhan Yuan, Kailai Yang, Zheheng Luo, Tianlin Zhang, Zhiwei Liu, Guojun Xiong, Zhiyang Deng, Yuechen Jiang, Zhiyuan Yao, Haohang Li, Yangyang Yu, Gang Hu, Jiajia Huang, Xiao-Yang Liu, Alejandro Lopez-Lira, Benyou Wang, Yanzhao Lai, Hao Wang, Min Peng, Sophia Ananiadou, Jimin Huang
arXiv:2402.12659 · 2026-07-27 공개 · arXiv · PDF
retrieval-augmented-generation forecasting gpt-4 information-extraction text-summarization risk-management financial-llm stock-trading
Abstract
LLMs have transformed NLP and shown promise in various fields, yet their potential in finance is underexplored due to a lack of comprehensive evaluation benchmarks, the rapid development of LLMs, and the complexity of financial tasks. In this paper, we introduce FinBen, the first extensive open-source evaluation benchmark, including 36 datasets spanning 24 financial tasks, covering seven critical aspects: information extraction (IE), textual analysis, question answering (QA), text generation, risk management, forecasting, and decision-making. FinBen offers several key innovations: a broader range of tasks and datasets, the first evaluation of stock trading, novel agent and Retrieval-Augmented Generation (RAG) evaluation, and three novel open-source evaluation datasets for text summarization, question answering, and stock trading. Our evaluation of 15 representative LLMs, including GPT-4, ChatGPT, and the latest Gemini, reveals several key findings: While LLMs excel in IE and textual analysis, they struggle with advanced reasoning and complex tasks like text generation and forecasting. GPT-4 excels in IE and stock trading, while Gemini is better at text generation and forecasting. Instruction-tuned LLMs improve textual analysis but offer limited benefits for complex tasks such as QA. FinBen has been used to host the first financial LLMs shared task at the FinNLP-AgentScen workshop during IJCAI-2024, attracting 12 teams. Their novel solutions outperformed GPT-4, showcasing FinBen's potential to drive innovation in financial LLMs. All datasets, results, and codes are released for the research community: https://github.com/The-FinAI/PIXIU.
한국어 요약
한 줄 요약
FinBen은 금융 분야 대형 언어 모델 평가를 위한 최초의 종합적 오픈소스 벤치마크로, 24개 금융 태스크와 36개 데이터셋을 포함한다.
핵심 기여도
- FinBen은 금융 분야 LLM 평가를 위한 최초의 종합적 오픈소스 벤치마크로, 24개 태스크, 36개 데이터셋을 포함.
- 7개 주요 영역(정보 추출, 텍스트 분석, QA, 텍스트 생성, 리스크 관리, 예측, 의사결정)으로 체계화.
- 주식 거래 평가, 에이전트 기반 평가, RAG 기반 평가 등 새로운 평가 전략 도입.
- 3개 새로운 오픈소스 평가 데이터셋(텍스트 요약, QA, 주식 거래) 제안.
핵심 아이디어
FinBen은 금융 분야에서의 LLM 활용을 체계적으로 평가하기 위해 설계된 벤치마크로, 기존 금융 NLP 평가가 주로 정보 추출과 QA에 집중한 반면, FinBen은 텍스트 생성, 예측, 의사결정 등 복잡한 태스크를 포함하여 평가 범위를 확장했다. 특히, 주식 거래와 같은 실시간 의사결정 태스크를 처음으로 평가에 포함시켰으며, Retrieval-Augmented Generation(RAG)과 에이전트 기반 평가를 도입함으로써 LLM이 실제 금융 환경에서 정보를 검색하고 활용하는 능력을 평가할 수 있도록 했다. FinBen은 24개 태스크와 36개 데이터셋을 통해 LLM의 종합적 성능을 측정하며, 15개 대표 모델(예: GPT-4, Gemini, LLaMA2-70B)을 평가하여 각 모델의 강점과 한계를 명확히 파악했다.
기술적 접근법
FinBen은 24개 금융 태스크를 7개 주요 영역(IE, TA, QA, TG, RM, FO, DM)으로 분류하고, 각 영역에 맞는 36개의 데이터셋을 구성하여 평가를 수행했다. 평가 전략으로는 기존의 단순 QA 평가를 넘어, RAG 기반 평가와 에이전트 기반 평가를 도입하여 LLM이 실제 금융 환경에서 정보를 검색하고 의사결정을 내리는 능력을 평가했다. 평가 대상 모델은 GPT-4, Gemini, LLaMA2-70B, ChatGLM3-6B 등 15개의 대표적인 일반 및 금융 전용 LLM이며, 최대 생성 토큰 수는 1024, 배치 크기는 20,000으로 설정하여 NVIDIA A100 80G GPU 16대에서 약 600시간 동안 실험을 수행했다.
주요 결과
- GPT-4는 정보 추출(IE)과 주식 거래(Stock Trading)에서 뛰어난 성능을 보였으며, QA 태스크에서도 우수한 결과를 기록.
- Gemini는 텍스트 생성(TG)과 예측(FO)에서 GPT-4보다 더 높은 성능을 나타냄.
- 지시어 튜닝된 LLM은 텍스트 분석(TA)에서 성능 향상이 있었으나, QA나 텍스트 생성과 같은 복잡한 태스크에서는 제한적.
- FinNLP-AgentScen 워크숍에서 12개 팀이 FinBen을 기반으로 개발한 모델이 GPT-4를 초과하는 성능을 보여, FinBen의 혁신 유도 가능성 입증.
의의 및 한계
FinBen은 금융 분야에서 LLM의 종합적 평가를 가능하게 하며, 기존 금융 NLP 평가가 단일 태스크에 집중한 한계를 극복했다. 특히, 주식 거래와 같은 실시간 의사결정 태스크를 포함한 평가 전략은 금융 LLM 연구에 새로운 방향을 제시한다. 그러나 데이터셋의 크기와 미국 시장 중심의 데이터는 모델의 일반화 능력에 영향을 줄 수 있으며, LLaMA 70B 모델만 평가한 점도 한계로 지적된다. 또한, 금융 정보의 오용 가능성에 대한 윤리적 고려가 필요하다는 점도 언급된다.
실용적 활용
FinBen은 금융 분석, 투자 전략 수립, 리스크 관리 등 다양한 금융 업무에서 LLM의 성능을 평가하는 데 활용될 수 있다. 또한, 금융 기관이 자체 개발한 LLM의 성능 검증 및 비교에 사용될 수 있으며, 학계에서는 금융 LLM 연구의 기준이 되는 벤치마크로 채택될 수 있다.