Business Arena: Benchmarking LLM Agents in a Realistic Marketplace

Yijun Pan, Yukun Lian, Kunyu Shi, Junbo Li, Hongwei Xue, Sicong Xie, Guannan Zhang, Xiaoying Xing

arXiv:2608.08621 · 2026-08-12 공개 · arXiv · PDF

llm-agents long-horizon agent-evaluation business-simulation profit-measurement skill-level-metrics action-attribution alibaba-data

Abstract

Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty, adapt to delayed outcomes in a changing market, and satisfy regulatory obligations before trading legally. Frontier LLM agents can increasingly complete complex workflows, yet business-related capabilities are rarely evaluated in existing agent benchmarks. We introduce Business Arena, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon. We ground the arena in real Alibaba.com sourcing data and market conditions calibrated from authoritative sources. Delayed and coupled consequences make individual business decisions difficult to judge, but their combined outcome is measurable through profit. Because profit alone cannot explain why an agent succeeds or fails, we compare agents with human-designed strategies to estimate available opportunity, use skill-level metrics to reveal underlying strengths and weaknesses, and trace realized gains and losses to the actions that produced them. We use mechanism ablations to establish that strong results reflect genuine business intelligence rather than neglect or simulator-specific shortcuts. We evaluate 15 frontier models and find a ninefold difference in mean final net worth. Even the best model falls behind human-designed strategies, indicating that business operation remains challenging for LLM agents. Skill-level analysis reveals operating styles, from margin-focused premium sellers to high-turnover wholesalers and customer-service specialists, while action-level attribution identifies the sourcing, pricing, and recovery decisions that create or destroy value. Together, Business Arena takes a first step toward a realistic and trustworthy testbed for evaluating end-to-end business agents.

한국어 요약

한 줄 요약

Business Arena는 AI 에이전트의 실제 비즈니스 운영 능력을 평가하는 벤치마크 환경으로, 15개 모델 간 최대 9배의 자산 차이를 기록했다.

핵심 기여도

핵심 아이디어

Business Arena는 AI 에이전트가 실제 비즈니스 환경에서 장기적으로 운영할 수 있는 능력을 평가하기 위해 설계되었다. 기존의 에이전트 벤치마크는 비즈니스 운영과 관련된 복합적인 도전 과제를 충분히 반영하지 못했다. 이 연구는 **시장 조사, 구매, 가격 설정, 고객 서비스, 컴플라이언스, 재무** 등 비즈니스 전 과정을 아우르는 **종단 간 루프**(end-to-end business loop)를 구현했다.

에이전트는 **부분적이고 모순적인 시장 신호**를 해석하고, **자본을 투자**한 뒤, **지연된 피드백**을 기반으로 전략을 조정해야 한다. 또한, **규제 준수**와 **지속적인 운영 의무**를 수행해야 하며, 이는 단순한 작업 수행을 넘어 **지속적인 판단과 적응**을 요구한다. 연구팀은 **기회 분석, 자본 배분, 마진 보존, 시장 학습** 등의 핵심 능력을 측정하기 위해 **스킬 레벨 메트릭**과 **액션 수준 속성 분석**(action-level attribution)을 도입했다.

기술적 접근법

주요 결과

의의 및 한계

Business Arena는 AI 에이전트의 **종단 간 비즈니스 운영 능력**을 평가하는 **실제적이고 신뢰할 수 있는 테스트베드**를 제공하며, **복잡한 시장 조건**과 **지속적인 운영 의무**를 반영한 첫 시도이다. **스킬 레벨 메트릭**과 **액션 수준 속성 분석**을 통해 모델의 **전략적 판단 능력**을 세부적으로 평가할 수 있어, 향후 모델 개선과 학습 데이터 구성에 유용하다.

하지만, **실제 비즈니스 환경의 모든 요소**를 완전히 재현하지 못하며, **규제 준수나 법적 책임**과 같은 **더 복잡한 요소**는 아직 포함되지 않았다. 또한, **모델의 실수로 인한 실제 자본 손실**을 방지하기 위해 **시뮬레이션 기반 평가**로 제한되어 있어, **실제 운영 환경에서의 일반화 가능성**은 추가 연구가 필요하다.

실용적 활용

Business Arena는 **AI 에이전트가 실제 비즈니스 운영에 얼마나 적합한지 평가**하는 데 활용될 수 있으며, **SaaS 기반 비즈니스 플랫폼, 자동화된 유통 시스템, 고객 서비스 자동화** 등 다양한 산업 분야에서 모델 성능을 검증하는 데 사용될 수 있다. 또한, **비즈니스 전략 개발, 마케팅 자동화, 재무 관리** 등 연구 분야에서도 유용한 평가 도구로 활용 가능하다.