JarvisHub:面向畫布原生多模態創意代理的開放式工具框架
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents
July 26, 2026
作者: Yunlong Lin, Zixu Lin, Zhaohu Xing, Biqiang Li, Chenxin Li, Haonan Wang, Haitao Wu, Hengyu Liu, Jianghai Chen, Kaituo Feng, Kaixin Li, Shawn Chen, Shijue Huang, Sixiang Chen, Tsung-Yi Ho, Wenxuan Huang, Xiangyan Liu, Xiaomeng Hu, Xuanhua He, Yan Sun, Yunqing Zhao, Zhiqin Yang, Zehan Wang, Zhengyang Tang, Tianyu Pang, Xiangyu Yue
cs.AI
摘要
创意AI正从单步资产生成向长周期多模态创作演进。尽管近期生成的模型能够合成高质量的图像、视频、音频片段、界面元素、故事板、幻灯片及其他创意资产,但现实世界的创意工作远不止孤立的提示输出交互。它涉及参考素材、草稿、备选方案、修改、失败尝试、版本关系、工具操作、评估信号及人类反馈,这些共同构成一个不断演变的项目状态。现有的基于提示、对话或节点的生成系统仅能部分支持这种状态,因为它们常丢弃中间上下文、依赖线性对话或需要手动指定工作流。近期商业化系统已显露出向智能体辅助创意生产转变的趋势,但其封闭架构使得研究智能体如何表征上下文、选择工具、修正制品、从失败中恢复以及随时间保持一致性变得困难。为弥补这一空白,我们提出JarvisHub——一款以画布为核心的创意智能体工具集,专为长周期多模态创作设计。JarvisHub将可编辑画布同时视为用户工作区、智能体的外部记忆、动作空间及共享项目状态,将多模态制品、依赖关系、版本及反馈表示为带类型的画布节点与链接。通过画布状态、协议桥梁与智能体运行时的三层架构,JarvisHub使智能体能够在可检查、可编辑的创作状态中行动。这一设计将创意智能体从孤立的工具使用推向可持续、可人为引导的创意自动化——智能体能够逐步规划、生成、修改并组织多模态项目,而用户始终能全程检查、引导并干预。
English
Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, audio clips, UI elements, storyboards, slides, and other creative assets, real-world creative work requires more than isolated prompt-output interactions. It involves references, drafts, alternatives, edits, failed attempts, version relations, tool actions, evaluation signals, and human feedback, which together form an evolving project state. Existing prompt-based, chat-based, and node-based generation systems only partially support this state, as they often discard intermediate context, rely on linear conversations, or require manually specified workflows. Recent commercial systems indicate a shift toward agent-assisted creative production, but their closed architectures make it difficult to study how agents represent context, choose tools, revise artifacts, recover from failures, and maintain consistency over time. To address this gap, we introduce JarvisHub, a canvas-native creative agent harness for long-horizon multimodal creation. JarvisHub treats an editable canvas as the user workspace, the agent's external memory, action space, and shared project state, representing multimodal artifacts, dependencies, versions, and feedback as typed canvas nodes and links. Through a three-layer architecture of canvas state, protocol bridge, and agent runtime, JarvisHub enables agents to act within an inspectable and editable creative state. This design moves creative agents beyond isolated tool use toward sustained, human-steerable creative automation, where agents can progressively plan, generate, revise, and organize multimodal projects while users remain able to inspect, guide, and intervene throughout the process.