ChatPaper.aiChatPaper

OmniAssistBench:面向全模態大語言模型的助理式互動基準

OmniAssistBench: Assistant-style Interaction Benchmark for Omni-LLMs

August 21, 2026
作者: Xianyun Sun, Chaoyou Fu, Zhengye Zhang, Feiyang Duan, Qingyuan Cao, Yonghui Niu, Sihang Yuan, Ge Zhang, Caifeng Shan
cs.AI

摘要

近期,全模態大型語言模型(Omni-LLMs)展現出作為即時影片助理的巨大潛力,能夠持續感知環境並引導使用者達成特定目標。與傳統被動式影片理解不同,互動式助理應主動結合視覺狀態、使用者目標與先驗知識,以提供有效的協助。評估此類能力極具挑戰性,因為模型不可預測的回應會動態改變使用者後續的行為,而靜態離線資料集無法涵蓋這種情境。為了解決此瓶頸,我們提出了 OmniAssistBench。為了解決同一使用者目標可透過多種方式達成而導致的互動路徑分歧問題,我們提供模型源自原始影片的預定義先驗資訊,要求其沿著完全相同的路徑引導使用者。由於真實互動影片相當稀少,我們透過逆向工程現有的網路影片來建構資料集。我們推導出合理的使用者目標,並將影片分割為多輪片段,以模擬連續互動。此嚴謹的流程耗費超過 1,000 小時的專家工時來建構資料集。結果顯示,專有的 Gemini-3-Pro 在滿分 100 分中獲得 66.4 分,而開源的 Qwen3-Omni-Instruct 則取得 51.2 分。儘管當前模型大致能理解使用者輸入,但它們經常提供錯誤或不完整的回答。具體而言,它們在處理視覺提示(如手勢)方面存在困難,在多輪互動中無法維持歷史脈絡,也無法延遲回應直到目標事件發生。結果顯示,在模型能成為可靠的助理之前,仍有相當大的改進空間。
English
Recent omni-modal large language models (Omni-LLMs) show great potential as real-time video assistants, which continuously perceive environments and guide users to achieve specific goals. Unlike traditional passive video understanding, interactive assistants should actively combine visual states, user goals, and prior knowledge to provide effective help. Evaluating this is rather challenging, as the model's unpredictable response dynamically changes the user's subsequent actions, which static offline datasets cannot accommodate. To address this bottleneck, we introduce OmniAssistBench. To solve the issue of diverging interaction paths where the same user goal can be achieved through various methods, we provide models with predefined priors derived from the source video, requiring them to guide users along the exact same routes. Since real interaction videos are rare, we construct the dataset by reverse-engineering existing Internet videos. We deduce logical user goals and segment the videos into multi-turn clips to simulate continuous interactions. This rigorous pipeline required over 1000 expert person-hours to build the dataset. Results show that the proprietary Gemini-3-Pro reaches 66.4 out of the max point of 100, while the open-source Qwen3-Omni-Instruct achieves 51.2. Although current models generally understand user inputs, they frequently provide incorrect or incomplete answers. Specifically, they struggle with visual prompts (e.g., hand gestures), fail to maintain historical context during multi-turn interactions, and fail to delay response until the target event. Results indicate substantial room for improvement before models can become reliable assistants.