MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

arXiv:2607.28956 · 2026-08-05 공개 · arXiv · PDF

llm-agents llm-evaluation benchmarking e-commerce long-term-coherence merchantbench order-simulation tool-interaction

Abstract

Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumulated evidence. Evaluating this capacity requires a persistent environment in which actions constrain future choices, feedback arrives at heterogeneous delays, and incoherent behavior produces measurable cumulative effects. Seller-side e-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash-Flow Management, and Mixed-Latency Feedback Adaptation. We introduce MerchantBench, a 365-day order-level simulation grounded in 98,843 real e-commerce product records and equipped with 26 tools for agent interaction. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions. We evaluate eight LLMs under two agent frameworks in 48 runs, each spanning 365 simulated days. Our results reveal a substantial gap between even the latest LLMs and human participants, with the best LLM configuration attaining only 27.3\% of the mean final net assets achieved by human participants.

한국어 요약

한 줄 요약

MerchantBench는 365일간 지속적인 온라인 판매 시뮬레이션을 통해 LLM 에이전트의 장기 일관성(Long-Term Coherence)을 평가하는 벤치마크이다.

핵심 기여도

핵심 아이디어

기존 LLM 평가가 단기적이고 명확한 성공 기준을 가진 작업에 집중하는 반면, 실제 온라인 판매는 장기적 일관성을 요구한다. 이에 따라, 연구팀은 온라인 판매 환경을 기반으로 **Long-Term Coherence**를 평가하는 새로운 벤치마크인 **MerchantBench**를 제안했다. 이 시뮬레이션은 상품 주문 수준에서의 **Upstream Supplier Events**와 **Downstream Order Outcomes**의 비동기적 피드백 구조를 반영하며, 에이전트가 이전 결정을 재검토하고 전략을 조정하도록 요구한다. 특히, **Product Sourcing**, **Pricing Control**, **Cash-Flow Management**, **Mixed-Latency Feedback Adaptation**이라는 4가지 상호 연관된 의사결정 과정을 통해 장기적 정책 유지와 수정 능력을 평가한다.

기술적 접근법

주요 결과

의의 및 한계

실용적 활용