JarvisHub:面向原生画布多模态创意智能体的开放框架
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents
July 26, 2026
作者: Yunlong Lin, Zixu Lin, Zhaohu Xing, Biqiang Li, Chenxin Li, Haonan Wang, Haitao Wu, Hengyu Liu, Jianghai Chen, Kaituo Feng, Kaixin Li, Shawn Chen, Shijue Huang, Sixiang Chen, Tsung-Yi Ho, Wenxuan Huang, Xiangyan Liu, Xiaomeng Hu, Xuanhua He, Yan Sun, Yunqing Zhao, Zhiqin Yang, Zehan Wang, Zhengyang Tang, Tianyu Pang, Xiangyu Yue
cs.AI
摘要
创意AI正从单步资产生成迈向长期多模态创作。尽管近期生成模型能够合成高质量图像、视频、音频片段、UI元素、故事板、演示文稿及其他创意资产,但真实世界的创意工作远不止于孤立的提示-输出交互。它涉及参考素材、草稿、备选方案、编辑修改、失败尝试、版本关系、工具操作、评估信号和人工反馈,这些共同构成一个不断演化的项目状态。现有的基于提示、对话或节点的生成系统仅能部分支持这种状态,因为它们常常丢弃中间上下文、依赖线性对话或需要手动指定工作流。近期商业系统已显示出向智能体辅助创意生产转变的趋势,但其封闭架构使得研究智能体如何表示上下文、选择工具、修订作品、从失败中恢复以及长期保持一致性变得困难。为弥补这一空白,我们提出JarvisHub——一个面向长期多模态创作的画布原生创意智能体框架。JarvisHub将可编辑画布视为用户工作空间、智能体外部记忆、动作空间和共享项目状态,将多模态作品、依赖关系、版本和反馈表示为带类型的画布节点与连接。通过画布状态、协议桥接和智能体运行时三层架构,JarvisHub使智能体能够在可检查、可编辑的创意状态中行动。这一设计将创意智能体从孤立的工具使用推向持续、可人为引导的创意自动化,智能体能够逐步规划、生成、修改和组织多模态项目,而用户在整个过程中可以检查、引导并随时介入。
English
Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, audio clips, UI elements, storyboards, slides, and other creative assets, real-world creative work requires more than isolated prompt-output interactions. It involves references, drafts, alternatives, edits, failed attempts, version relations, tool actions, evaluation signals, and human feedback, which together form an evolving project state. Existing prompt-based, chat-based, and node-based generation systems only partially support this state, as they often discard intermediate context, rely on linear conversations, or require manually specified workflows. Recent commercial systems indicate a shift toward agent-assisted creative production, but their closed architectures make it difficult to study how agents represent context, choose tools, revise artifacts, recover from failures, and maintain consistency over time. To address this gap, we introduce JarvisHub, a canvas-native creative agent harness for long-horizon multimodal creation. JarvisHub treats an editable canvas as the user workspace, the agent's external memory, action space, and shared project state, representing multimodal artifacts, dependencies, versions, and feedback as typed canvas nodes and links. Through a three-layer architecture of canvas state, protocol bridge, and agent runtime, JarvisHub enables agents to act within an inspectable and editable creative state. This design moves creative agents beyond isolated tool use toward sustained, human-steerable creative automation, where agents can progressively plan, generate, revise, and organize multimodal projects while users remain able to inspect, guide, and intervene throughout the process.