ChatPaper.aiChatPaper

JarvisHub: キャンバスネイティブなマルチモーダルクリエイティブエージェントのためのオープンハーネス

JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents

July 26, 2026
著者: Yunlong Lin, Zixu Lin, Zhaohu Xing, Biqiang Li, Chenxin Li, Haonan Wang, Haitao Wu, Hengyu Liu, Jianghai Chen, Kaituo Feng, Kaixin Li, Shawn Chen, Shijue Huang, Sixiang Chen, Tsung-Yi Ho, Wenxuan Huang, Xiangyan Liu, Xiaomeng Hu, Xuanhua He, Yan Sun, Yunqing Zhao, Zhiqin Yang, Zehan Wang, Zhengyang Tang, Tianyu Pang, Xiangyu Yue
cs.AI

要旨

クリエイティブAIは、単一工程のアセット生成から長期的なマルチモーダル制作へと移行しつつある。近年の生成モデルは高品質な画像、動画、音声クリップ、UI要素、ストーリーボード、スライド、その他のクリエイティブアセットを合成できるが、現実のクリエイティブ作業は孤立したプロンプトと出力のやりとりだけでは完結しない。そこには参照、下書き、代替案、編集、失敗例、バージョン間の関係、ツール操作、評価シグナル、人間からのフィードバックが含まれ、これらが全体として進化するプロジェクト状態を形成する。既存のプロンプトベース、チャットベース、ノードベースの生成システムは、中間コンテキストを破棄したり、直線的な会話に依存したり、手動でワークフローを指定する必要があったりと、この状態を部分的にしかサポートしていない。最近の商用システムはエージェント支援型のクリエイティブ制作へのシフトを示しているが、その閉じたアーキテクチャのため、エージェントがどのようにコンテキストを表現し、ツールを選択し、アーティファクトを修正し、障害から回復し、時間をかけて一貫性を維持するかを研究することが困難である。このギャップを埋めるために、我々はJarvisHubを紹介する。これは長期的なマルチモーダル制作のためのキャンバスネイティブなクリエイティブエージェントハーネスである。JarvisHubは編集可能なキャンバスをユーザーのワークスペース、エージェントの外部記憶、行動空間、共有プロジェクト状態として扱い、マルチモーダルアーティファクト、依存関係、バージョン、フィードバックを型付きのキャンバスノードとリンクとして表現する。キャンバス状態、プロトコルブリッジ、エージェントランタイムの3層アーキテクチャにより、JarvisHubはエージェントが検査可能かつ編集可能なクリエイティブ状態の中で行動することを可能にする。この設計はクリエイティブエージェントを孤立したツール使用から、持続的かつ人間が操縦可能なクリエイティブ自動化へと押し上げる。ここでエージェントはマルチモーダルプロジェクトを段階的に計画、生成、修正、整理できる一方、ユーザーはプロセス全体を通じて検査、指導、介入することができる。
English
Creative AI is moving from single-step asset generation toward long-horizon multimodal production. Although recent generative models can synthesize high-quality images, videos, audio clips, UI elements, storyboards, slides, and other creative assets, real-world creative work requires more than isolated prompt-output interactions. It involves references, drafts, alternatives, edits, failed attempts, version relations, tool actions, evaluation signals, and human feedback, which together form an evolving project state. Existing prompt-based, chat-based, and node-based generation systems only partially support this state, as they often discard intermediate context, rely on linear conversations, or require manually specified workflows. Recent commercial systems indicate a shift toward agent-assisted creative production, but their closed architectures make it difficult to study how agents represent context, choose tools, revise artifacts, recover from failures, and maintain consistency over time. To address this gap, we introduce JarvisHub, a canvas-native creative agent harness for long-horizon multimodal creation. JarvisHub treats an editable canvas as the user workspace, the agent's external memory, action space, and shared project state, representing multimodal artifacts, dependencies, versions, and feedback as typed canvas nodes and links. Through a three-layer architecture of canvas state, protocol bridge, and agent runtime, JarvisHub enables agents to act within an inspectable and editable creative state. This design moves creative agents beyond isolated tool use toward sustained, human-steerable creative automation, where agents can progressively plan, generate, revise, and organize multimodal projects while users remain able to inspect, guide, and intervene throughout the process.