ChatPaper.aiChatPaper

Vinci2: 継続的な一人称視点動画における先回り支援

Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos

July 13, 2026
著者: Gong Sitong, Tianyu Yan, Caixin Kang, Bo Zheng, Xiang Ruan, Huchuan Lu, Kaipeng Zhang, Yoichi Sato, Yifei Huang
cs.AI

要旨

インテリジェントアシスタントは、いつ能動的に発言すべきか。連続的な自己中心視点映像は、豊かで進化する文脈を提供し、単なる反応型ではなく能動型の新たな支援を可能にする。しかし、既存のアプローチでは、ユーザの問い合わせを消極的に待つか、検出されたあらゆるイベントを応答すべきものとして扱い、ユーザの履歴や現在の活動、支援が実際に望まれるかどうかを考慮しない。我々は能動的支援を、状況に依存した意思決定問題として再定義する。すなわち、エージェントは何が起きているかを認識するだけでなく、蓄積された時間的文脈に基づいて推論し、いつ、そして介入すべきか否かを判断しなければならない。この目的のため、我々はVinci2を提案する。これは、オンデバイスアシスタントVinciを反応型応答から能動性へと進化させる、能動的自己中心視点支援システムである。評価面では、連続的な自己中心視点映像における能動的支援のための初の大規模ベンチマークであるEgoServeを提示する。EgoServeは、即時的な安全警告から長期にわたる習慣指導まで、4つの時間的記憶範囲(temporal memory horizons)に沿って構成された3,000以上のサービスインスタンスを、10のサービスカテゴリにわたって含む。モデリング面では、学習不要の記憶拡張エージェントEgoMemoを提案する。これは、マルチスケールの時間的要約、意味的知識グラフ、視覚的埋め込みアーカイブという、相補的な三つの記憶表現を維持する。各タイムステップにおいて、EgoMemoは検索拡張推論(retrieval-augmented reasoning)を実行して支援が適切かどうかを判断し、適切であれば文脈に基づいた応答を生成する。実験により、EgoMemoは既存の自己中心視点ベンチマークにおいて競争力を保ちつつ、EgoServe上で強力なベースラインを確立することを示す。我々のベンチマークとコードはhttps://sitonggong.github.io/EgoServe-page/{Vinci2}で公開されている。
English
When should an intelligent assistant speak up without being asked? Continuous egocentric video offers rich, evolving context that enables a new form of assistance: one that is proactive rather than merely reactive. Yet existing approaches either wait passively for user queries or treat every detected event as requiring a response, without considering the user's history, current activity, or whether assistance would actually be welcome. We reframe proactive assistance as a context-dependent decision problem: the agent must not only perceive what is happening, but reason over accumulated temporal context to determine when and whether to intervene. To this end, we present Vinci2, a proactive egocentric assistance system that advances the on-device assistant Vinci from reactive response toward proactivity. On the evaluation side, we present EgoServe, the first large-scale benchmark for proactive assistance in continuous egocentric video. EgoServe comprises over 3,000 service instances organized along 4 temporal memory horizons, ranging from immediate safety alerts to long-term habit coaching, across 10 service categories. On the modeling side, we propose EgoMemo, a training-free, memory-augmented agent that maintains three complementary memory representations: multi-scale temporal summaries, a semantic knowledge graph, and visual embedding archives. At each timestep, EgoMemo performs retrieval-augmented reasoning to determine whether assistance is warranted and, if so, produces contextually grounded responses. Experiments demonstrate that EgoMemo establishes strong baselines on EgoServe while remaining competitive on existing egocentric benchmarks. Our benchmark and code are publicly available at https://sitonggong.github.io/EgoServe-page/{Vinci2}.