SimuVerity: Benchmarking Agents for Engineering-Grade Simulink Model Generation

Ruiqi Zhang, Jiahao Wang, Mingxuan Li, Haichen Luo, Chaoting Wang, Guoyu Mou, Keyu Lai, Hanchao Lv, Jiaxu Wang, Yibo Zheng, Aijun Yang, Xiaohua Wang

arXiv:2610.02304 · 2026-10-05 공개 · arXiv · PDF

agent-evaluation simulink model-generation engineering-verification executable-models simulation-scenarios mechanistic-fidelity control-integrity

Abstract

Existing Simulink benchmarks mainly evaluate whether generated models compile, execute, or resemble a reference model. These criteria do not establish whether a model satisfies its engineering requirements. We introduce SimuVerity, a benchmark of 101 text-to-executable Simulink model-generation tasks across ten engineering domains. For each task, executable-system profiles ground the engineering specification and four families of native simulation scenarios. A hierarchical evaluator first checks artifact delivery, native executability, and engineering qualification, then scores qualified models across six dimensions covering accuracy, output quality, mechanistic fidelity, control and causal integrity, operating-domain robustness, and dynamic response. We evaluate six agent systems with SimuVerity. The best system achieves an overall score of only 42.86. The results show that structural similarity is a poor proxy for engineering performance: capability bottlenecks arise both in producing qualified implementations and in satisfying multidimensional requirements after qualification. Meanwhile, some high-scoring models still exhibit severe visual-layout disorder. SimuVerity provides a systematic basis for assessing agents' engineering capabilities and diagnosing failures in executable Simulink model generation.

한국어 요약

한 줄 요약

SimuVerity는 10개 공학 분야에서 101개의 텍스트-실행가능 Simulink 모델 생성 작업을 평가하는 벤치마크로, 최고 시스템의 종합 점수는 42.86에 불과하다.

핵심 기여도

핵심 아이디어

기존 Simulink 벤치마크는 생성된 모델이 컴파일되거나 실행되거나 참조 모델과 유사한지 여부만 평가하며, 공학적 요구사항을 충족하는지 확인하지 못한다. SimuVerity는 각 작업에 대해 실행 가능한 시스템 프로파일과 네 가지 시뮬레이션 시나리오를 통해 공학적 명세를 구체화하고, A/Q/M/C/R/D 6차원 평가 체계를 도입하여 모델의 공학적 성능을 다차원적으로 평가한다. 이는 단순한 구조적 유사성 대신 실제 공학적 요구사항을 중심으로 평가한다는 점에서 혁신적이다.

기술적 접근법

주요 결과

의의 및 한계

SimuVerity는 공학적 요구사항을 중심으로 Simulink 모델 생성 에이전트를 평가하는 체계적인 기반을 제공하며, 단순한 실행 가능성이나 구조적 유사성에 의존하는 기존 평가 방식의 한계를 보완한다. 그러나 MATLAB/Simulink에 의존하는 점에서 평가 속도가 느리고, 확장성에 제약이 있다. 또한, 일부 고점수 모델이 시각적 레이아웃 문제를 보이는 점은 공학적 성능과 사용자 경험 간 균형이 필요함을 시사한다.

실용적 활용

SimuVerity는 자동차, 항공, 제어 시스템 등 다양한 공학 분야에서 모델 기반 설계(MBD)를 수행하는 연구자와 엔지니어에게 유용한 평가 도구로 활용될 수 있다. 특히, LLM 기반 에이전트가 생성한 Simulink 모델의 공학적 신뢰도를 검증하고, 개선 방향을 도출하는 데 기여할 수 있다.