Vinci2: Proactieve ondersteuning bieden in continue egocentrische video's
Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos
July 13, 2026
Auteurs: Gong Sitong, Tianyu Yan, Caixin Kang, Bo Zheng, Xiang Ruan, Huchuan Lu, Kaipeng Zhang, Yoichi Sato, Yifei Huang
cs.AI
Samenvatting
Wanneer moet een intelligente assistent uit zichzelf spreken zonder dat erom gevraagd wordt? Continue egocentrische video biedt een rijke, evoluerende context die een nieuwe vorm van assistentie mogelijk maakt: een die proactief is in plaats van louter reactief. Toch wachten bestaande benaderingen passief op gebruikersvragen of behandelen ze elke gedetecteerde gebeurtenis alsof deze een reactie vereist, zonder rekening te houden met de geschiedenis van de gebruiker, diens huidige activiteit, of assistentie daadwerkelijk welkom zou zijn. Wij herformuleren proactieve assistentie als een contextafhankelijk beslissingsprobleem: de agent moet niet alleen waarnemen wat er gebeurt, maar ook redeneren over de opgebouwde temporele context om te bepalen wanneer en of ingrijpen nodig is. Hiertoe presenteren we Vinci2, een proactief egocentrisch assistentiesysteem dat de apparaatgebonden assistent Vinci van reactieve respons naar proactiviteit brengt. Aan de evaluatiekant presenteren we EgoServe, de eerste grootschalige benchmark voor proactieve assistentie in continue egocentrische video. EgoServe omvat meer dan 3.000 service-instanties, georganiseerd langs 4 temporele geheugenhorizons, variërend van onmiddellijke veiligheidsmeldingen tot langdurige gewoontecoaching, verdeeld over 10 servicecategorieën. Aan de modelleerkant stellen we EgoMemo voor, een trainingsvrije, geheugenversterkte agent die drie complementaire geheugenrepresentaties onderhoudt: meerschalige temporele samenvattingen, een semantische kennisgraaf en archieven van visuele inbeddingen. Bij elke tijdstap voert EgoMemo ophaalversterkte redenering uit om te bepalen of assistentie gerechtvaardigd is en, zo ja, produceert het contextueel gefundeerde reacties. Experimenten tonen aan dat EgoMemo sterke basislijnen vestigt op EgoServe, terwijl het concurrerend blijft op bestaande egocentrische benchmarks. Onze benchmark en code zijn publiekelijk beschikbaar op https://sitonggong.github.io/EgoServe-page/{Vinci2}.
English
When should an intelligent assistant speak up without being asked? Continuous egocentric video offers rich, evolving context that enables a new form of assistance: one that is proactive rather than merely reactive. Yet existing approaches either wait passively for user queries or treat every detected event as requiring a response, without considering the user's history, current activity, or whether assistance would actually be welcome. We reframe proactive assistance as a context-dependent decision problem: the agent must not only perceive what is happening, but reason over accumulated temporal context to determine when and whether to intervene. To this end, we present Vinci2, a proactive egocentric assistance system that advances the on-device assistant Vinci from reactive response toward proactivity. On the evaluation side, we present EgoServe, the first large-scale benchmark for proactive assistance in continuous egocentric video. EgoServe comprises over 3,000 service instances organized along 4 temporal memory horizons, ranging from immediate safety alerts to long-term habit coaching, across 10 service categories. On the modeling side, we propose EgoMemo, a training-free, memory-augmented agent that maintains three complementary memory representations: multi-scale temporal summaries, a semantic knowledge graph, and visual embedding archives. At each timestep, EgoMemo performs retrieval-augmented reasoning to determine whether assistance is warranted and, if so, produces contextually grounded responses. Experiments demonstrate that EgoMemo establishes strong baselines on EgoServe while remaining competitive on existing egocentric benchmarks. Our benchmark and code are publicly available at https://sitonggong.github.io/EgoServe-page/{Vinci2}.