JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents
Yunlong Lin, Zixu Lin, Zhaohu Xing, Biqiang Li, Chenxin Li, Haonan Wang, Haitao Wu, Hengyu Liu, Jianghai Chen, Kaituo Feng, Kaixin Li, Shawn Chen, Shijue Huang, Sixiang Chen, Tsung-Yi Ho, Wenxuan Huang, Xiangyan Liu, Xiaomeng Hu, Xuanhua He, Yan Sun, Yunqing Zhao, Zhiqin Yang, Zehan Wang, Zhengyang Tang, Tianyu Pang, Xiangyu Yue
arXiv:2607.23588 · 2026-07-28 공개 · arXiv · PDF
tool-use agent-runtime canvas-native multimodal-creation creative-agents editable-canvas project-state version-control
Abstract
Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, audio clips, UI elements, storyboards, slides, and other creative assets, real-world creative work requires more than isolated prompt-output interactions. It involves references, drafts, alternatives, edits, failed attempts, version relations, tool actions, evaluation signals, and human feedback, which together form an evolving project state. Existing prompt-based, chat-based, and node-based generation systems only partially support this state, as they often discard intermediate context, rely on linear conversations, or require manually specified workflows. Recent commercial systems indicate a shift toward agent-assisted creative production, but their closed architectures make it difficult to study how agents represent context, choose tools, revise artifacts, recover from failures, and maintain consistency over time. To address this gap, we introduce JarvisHub, a canvas-native creative agent harness for long-horizon multimodal creation. JarvisHub treats an editable canvas as the user workspace, the agent's external memory, action space, and shared project state, representing multimodal artifacts, dependencies, versions, and feedback as typed canvas nodes and links. Through a three-layer architecture of canvas state, protocol bridge, and agent runtime, JarvisHub enables agents to act within an inspectable and editable creative state. This design moves creative agents beyond isolated tool use toward sustained, human-steerable creative automation, where agents can progressively plan, generate, revise, and organize multimodal projects while users remain able to inspect, guide, and intervene throughout the process.
한국어 요약
한 줄 요약
JarvisHub는 캔버스 기반의 창의적 에이전트 허브로, 장기적 멀티모달 창작 과정을 관리하고 추적 가능하게 만든다.
핵심 기여도
- 창의적 프로젝트를 **편집 가능한 캔버스 그래프**로 정식화하여 에이전트가 접근 가능하게 함.
- **Canvas State, Protocol Bridge, Agent Runtime**의 3단계 아키텍처로 프로젝트 상태를 관리 및 추적.
- **Model Context Protocol (MCP)** 기반 확장성을 지원하는 툴 패밀리 구현.
- 장기적 창작 과정에서 **트래젝토리 기록**을 통해 에이전트 행동 분석 가능하게 함.
핵심 아이디어
기존 창의적 AI는 단일 프롬프트-출력 방식에 머무르며, 실제 창작 과정에서 필요한 **중간 자료, 피드백, 버전 관계** 등을 제대로 반영하지 못한다. JarvisHub는 이러한 문제를 해결하기 위해 **캔버스를 공유 프로젝트 상태**로 사용하는 새로운 접근법을 제시한다. 사용자와 에이전트는 캔버스를 통해 **노드와 링크로 표현된 멀티모달 자산**을 생성, 수정, 검토할 수 있다. 이는 단순한 텍스트 기반 대화나 노드 기반 워크플로우와 달리, **공간적 레이아웃, 버전 관계, 피드백** 등을 포함한 진정한 프로젝트 상태를 유지할 수 있게 한다. 핵심 아이디어는 캔버스가 단순한 UI가 아니라 **에이전트가 읽고 쓸 수 있는 외부 메모리**라는 점이다.
기술적 접근법
- **Canvas State Layer**: 노드, 속성, 레이아웃, 버전, 의존성 링크를 저장.
- **Protocol Bridge**: 에이전트의 캔버스 읽기/쓰기 권한, 연산 형식, 검증, 로깅을 관리.
- **Agent Runtime**: 허가된 액션 선택, 툴 호출, 상태 동기화, 트래젝토리 기록.
- **툴 패밀리**: 캔버스 편집, 미디어 생성, 네이티브 실행, 복구, MCP 기반 확장 지원.
- **MCP (Model Context Protocol)**: 에이전트와 생성 모델 간의 컨텍스트 전달을 표준화.
주요 결과
- **장기적 멀티모달 창작 작업**에서 JarvisHub는 **트래젝토리 기록**을 통해 에이전트 행동을 추적 가능하게 함.
- **대화형 웹 개발, 서사 미디어 생성** 등 고가치 작업에서 **반복적 자산 생성과 피드백 기반 개선**을 지원.
- **에이전트의 결정 과정**이 **시각적 캔버스 노드**로 표현되어 **투명성과 재사용성** 향상.
의의 및 한계
JarvisHub는 창의적 에이전트 연구에서 **프로젝트 상태 관리와 에이전트 행동 추적**이라는 핵심 문제를 해결하는 기반 허브를 제공한다. 특히, **캔버스 기반의 공유 상태 표현**은 기존의 단일 프롬프트 기반 시스템과 노드 기반 워크플로우 도구의 한계를 극복한다. 또한, **트래젝토리 데이터**는 향후 창의적 에이전트의 평가 및 훈련에 유용한 자료로 활용될 수 있다. 그러나 현재 실험은 **정성적 시연**에 머무르며, 완전한 벤치마크나 리더보드는 제공되지 않았다. 또한, **최종 자산 품질**은 외부 모델과 도구에 의존하며, **에이전트의 창의적 결정이 의미적으로 올바른지 여부**는 보장되지 않는다.
실용적 활용
JarvisHub는 **UI/UX 디자인, 스토리보드 제작, 마케팅 콘텐츠 생성** 등 멀티모달 창작이 필요한 산업에서 활용 가능하다. 특히, **에이전트가 반복적 수정과 피드백을 반영하는 장기적 프로젝트**에서 유용하며, **창의적 워크플로우의 자동화와 투명성**을 동시에 요구하는 연구 및 개발 환경에 적합하다.