ChatPaper.aiChatPaper

JarvisHub: 캔버스 네이티브 멀티모달 창작 에이전트를 위한 개방형 하네스

JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

July 26, 2026
저자: Yunlong Lin, Zixu Lin, Zhaohu Xing, Biqiang Li, Chenxin Li, Haonan Wang, Haitao Wu, Hengyu Liu, Jianghai Chen, Kaituo Feng, Kaixin Li, Shawn Chen, Shijue Huang, Sixiang Chen, Tsung-Yi Ho, Wenxuan Huang, Xiangyan Liu, Xiaomeng Hu, Xuanhua He, Yan Sun, Yunqing Zhao, Zhiqin Yang, Zehan Wang, Zhengyang Tang, Tianyu Pang, Xiangyu Yue
cs.AI

초록

창의적 AI는 단일 단계의 자산 생성에서 장기적 멀티모달 생산으로 진화하고 있다. 최근 생성 모델들은 고품질의 이미지, 비디오, 오디오 클립, UI 요소, 스토리보드, 슬라이드 및 기타 창의적 자산을 합성할 수 있지만, 실제 창작 작업은 고립된 프롬프트-출력 상호작용 이상을 요구한다. 창작 작업에는 참고자료, 초안, 대안, 수정, 실패한 시도, 버전 관계, 도구 동작, 평가 신호 및 인간의 피드백이 포함되며, 이 모든 것이 진화하는 프로젝트 상태를 형성한다. 기존의 프롬프트 기반, 채팅 기반, 노드 기반 생성 시스템은 중간 맥락을 폐기하거나 선형적 대화에 의존하거나 수동 작업 흐름을 요구하는 경우가 많아 이러한 상태를 부분적으로만 지원한다. 최근 상용 시스템들은 에이전트 지원 창작 생산으로의 전환을 시사하지만, 폐쇄적 아키텍처로 인해 에이전트가 맥락을 표현하고, 도구를 선택하며, 결과물을 수정하고, 실패에서 복구하며, 시간이 지남에 따라 일관성을 유지하는 방식을 연구하기 어렵다. 이러한 격차를 해소하기 위해 우리는 JarvisHub를 소개한다. 이는 장기적 멀티모달 창작을 위한 캔버스 네이티브 창의적 에이전트 도구이다. JarvisHub는 편집 가능한 캔버스를 사용자 작업 공간, 에이전트의 외부 메모리, 행동 공간 및 공유 프로젝트 상태로 간주하며, 멀티모달 결과물, 의존성, 버전 및 피드백을 타입화된 캔버스 노드와 링크로 표현한다. 캔버스 상태, 프로토콜 브리지, 에이전트 런타임의 3계층 아키텍처를 통해 JarvisHub는 에이전트가 검사 가능하고 편집 가능한 창의적 상태 내에서 작동할 수 있게 한다. 이 설계는 창의적 에이전트를 고립된 도구 사용에서 벗어나 지속적이고 인간이 제어 가능한 창작 자동화로 이끌며, 에이전트가 점진적으로 계획하고, 생성하고, 수정하고, 멀티모달 프로젝트를 구성하는 동안 사용자는 전 과정을 검사하고 안내하며 개입할 수 있게 한다.
English
Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, audio clips, UI elements, storyboards, slides, and other creative assets, real-world creative work requires more than isolated prompt-output interactions. It involves references, drafts, alternatives, edits, failed attempts, version relations, tool actions, evaluation signals, and human feedback, which together form an evolving project state. Existing prompt-based, chat-based, and node-based generation systems only partially support this state, as they often discard intermediate context, rely on linear conversations, or require manually specified workflows. Recent commercial systems indicate a shift toward agent-assisted creative production, but their closed architectures make it difficult to study how agents represent context, choose tools, revise artifacts, recover from failures, and maintain consistency over time. To address this gap, we introduce JarvisHub, a canvas-native creative agent harness for long-horizon multimodal creation. JarvisHub treats an editable canvas as the user workspace, the agent's external memory, action space, and shared project state, representing multimodal artifacts, dependencies, versions, and feedback as typed canvas nodes and links. Through a three-layer architecture of canvas state, protocol bridge, and agent runtime, JarvisHub enables agents to act within an inspectable and editable creative state. This design moves creative agents beyond isolated tool use toward sustained, human-steerable creative automation, where agents can progressively plan, generate, revise, and organize multimodal projects while users remain able to inspect, guide, and intervene throughout the process.