Vinci2:在連續第一人稱影片中提供主動協助
Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos
July 13, 2026
作者: Gong Sitong, Tianyu Yan, Caixin Kang, Bo Zheng, Xiang Ruan, Huchuan Lu, Kaipeng Zhang, Yoichi Sato, Yifei Huang
cs.AI
摘要
智慧助理何時應在未被詢問時主動發聲?連續的自我中心影片提供了豐富且不斷演變的情境,從而實現一種新型的協助方式:主動式而非僅是被動反應式的協助。然而,現有方法要麼被動等待使用者提問,要麼將每一個偵測到的事件都視為需要回應,而未考慮使用者的歷史記錄、當前活動,或協助本身是否真的受歡迎。我們將主動協助重新定義為一個依賴情境的決策問題:代理者不僅須感知當下發生的事,更須運用累積的時間脈絡進行推理,以決定何時、甚至是否應介入。為此,我們提出Vinci2,一套主動式自我中心協助系統,將裝置上的助理Vinci從被動反應推向主動服務。在評測方面,我們提出EgoServe,首個針對連續自我中心影片中主動協助的大型基準。EgoServe包含超過3,000個服務實例,依據4種時間記憶範圍(從即時安全警報到長期習慣指導)組織,涵蓋10個服務類別。在建模方面,我們提出EgoMemo,一種免訓練、記憶增強的代理者,維護三種互補的記憶表徵:多尺度時間摘要、語意知識圖譜,以及視覺嵌入檔案。在每個時間步,EgoMemo執行檢索增強推理,判斷協助是否必要,並在必要時產出基於情境的回應。實驗顯示,EgoMemo在EgoServe上建立了強勁的基準表現,同時在現有自我中心基準上也具競爭力。我們的基準與程式碼已公開於 https://sitonggong.github.io/EgoServe-page/{Vinci2}。
English
When should an intelligent assistant speak up without being asked? Continuous egocentric video offers rich, evolving context that enables a new form of assistance: one that is proactive rather than merely reactive. Yet existing approaches either wait passively for user queries or treat every detected event as requiring a response, without considering the user's history, current activity, or whether assistance would actually be welcome. We reframe proactive assistance as a context-dependent decision problem: the agent must not only perceive what is happening, but reason over accumulated temporal context to determine when and whether to intervene. To this end, we present Vinci2, a proactive egocentric assistance system that advances the on-device assistant Vinci from reactive response toward proactivity. On the evaluation side, we present EgoServe, the first large-scale benchmark for proactive assistance in continuous egocentric video. EgoServe comprises over 3,000 service instances organized along 4 temporal memory horizons, ranging from immediate safety alerts to long-term habit coaching, across 10 service categories. On the modeling side, we propose EgoMemo, a training-free, memory-augmented agent that maintains three complementary memory representations: multi-scale temporal summaries, a semantic knowledge graph, and visual embedding archives. At each timestep, EgoMemo performs retrieval-augmented reasoning to determine whether assistance is warranted and, if so, produces contextually grounded responses. Experiments demonstrate that EgoMemo establishes strong baselines on EgoServe while remaining competitive on existing egocentric benchmarks. Our benchmark and code are publicly available at https://sitonggong.github.io/EgoServe-page/{Vinci2}.