AVA-Encoder: エージェントネイティブなビデオ表現学習に向けて
AVA-Encoder: Towards Agent-Native Video Representation Learning
August 12, 2026
著者: Chuyue Li, Jinpeng Yu, Haozhe Wang, Tian Xueyun, Zhijing Zhang, Bingnan Li, Shuqi Gu, Kan Ren, Jiaming Liu, Ruihua Hua
cs.AI
要旨
クリエイティブエージェントは、高品質な人間の映画から学習する効果的な手段を依然として欠いており、映画品質の映像を生成する能力が制限されている。重要な課題は、映画の内容に忠実でありながら、エージェントの推論と操作に直接利用可能な、構造化された映像表現が存在しないことである。この課題に対処するため、我々はエージェント的オートエンコーディング(Agentic Auto-Encoding)を通じてエージェント原生の映像表現を学習するフレームワークであるAgentic Video Auto-Encoder(AVA-Encoder)を提案する。
AVA-Encoderは、映像を知識グラフ(KG)表現に変換し、それを再び映像に再構成する。その階層構造と状態ノードは構造化テキストを格納し、リンクされたアセット層は生成された画像、音声、映像を保持する。型付きエッジは、これらのテキスト記述とアセット間の関係を、エージェントが容易に理解、検索、編集できる形式で保存する。映像再構成の差分はテキスト勾配最適化フレームワークを駆動し、評価フィードバックを自然言語による更新方向として表現することで、外側ループにおけるデータ非依存エンコーディングポリシー疑似訓練(Data-Independent Encoding Policy Pseudo-Training)と、テスト時内側ループにおける任意のデータ依存型KG表現洗練(Data-Dependent KG Representation Refinement)を実現する。
大規模な実験により、AVA-Encoderは最強の外部ベースラインを20.7パーセントポイント上回る改善を示す。制御されたポリシーのみの設定では、その疑似訓練されたショットレベルのAgentic Video Encoderポリシーは、慎重に人手で調整されたポリシーをも上回りながら、システムプロンプトのトークンを74.3%削減する。我々は、完全なAVA-Encoderフレームワーク、信頼性の高いエージェント的映像再構成ベンチマーク、および高品質な映画KG表現の最初のデータセットを公開する。
English
Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured video representation that is both faithful to film content and directly usable for agentic reasoning and manipulation. To address the challenge, we propose the Agentic Video Auto-Encoder (AVA-Encoder), a framework for learning agent-native video representations via agentic auto-encoding.
AVA-Encoder transforms a video into a knowledge graph (KG) representation and then reconstructs it back into video. Its hierarchy and state nodes store structured text, while a linked asset layer holds generated images, audio, and video. Typed edges preserve the relations between these text descriptions and assets in a form that agents can easily understand, query, and edit. The video reconstruction differences drive a textual-gradient optimization framework, which expresses evaluation feedback as natural-language update directions for Data-Independent Encoding Policy Pseudo-Training in the outer loop and optional Data-Dependent KG Representation Refinement in the test-time inner loop.
Extensive experiments show that AVA-Encoder improves by 20.7 percentage points over the strongest external baseline. In the controlled policy-only setting, its pseudo-trained shot-level Agentic Video Encoder policy also outperforms a carefully human-tuned policy while using 74.3% fewer system-prompt tokens. We release the complete AVA-Encoder framework, a reliable agentic video reconstruction benchmark, and the first dataset of high-quality film KG representations.