VaseMuseum:古希臘陶器數位智慧博物館
VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery
July 7, 2026
作者: Jiazi Wang, Nonghai Zhang, Qiushi Xie, Zeyu Zhang, Yufeng Chen, Yang Zhao, Ling Shao, Hao Tang
cs.AI
摘要
視覺語言模型(VLMs)透過將3D數位化與自然語言文物探索相結合,使互動式數位博物館日益可行。然而,在古希臘陶器等文化遺產領域中,可靠的VLM輔助仍面臨兩項挑戰。首先,開放式詮釋需要將細粒度的2D/3D視覺證據與專業策展知識相結合,但檢索過程可能引入薄弱來源與不可驗證的參考資料。其次,當可用證據不完整、有雜訊或模稜兩可時,VLM常產生看似可信卻缺乏依據的回應,而非校準後的不確定性表達。為解決這些挑戰,我們提出VaseMuseum——一個專為古希臘陶器智慧數位博物館設計的輕量化模組化多模態智能體框架。VaseMuseum結合互動式虛擬博物館與VaseAgent,後者透過多模態感知、3D感知推理、外部知識檢索與推理時可靠性控制,同時支援2D影像與3D文物。具體而言,VaseAgent從權威網站與博物館知識來源檢索證據,並透過來源層級控制,在生成前選取多樣化且可驗證的證據。同時,回應層級控制針對生成的主張與證據池進行比對,在證據不足或衝突時鼓勵中性、受證據約束的回應。此外,我們採用無需訓練的GRPO風格選擇機制,在不更新VLM主幹的前提下,傾向選取具有效參考與校準信心的回應。在真實數位博物館模擬實驗中,VaseMuseum相較於啟用搜尋的VLM基準,能提升引用有效性、減少知識密集型查詢的幻覺,並在模糊情境下產生更中性的回應。
English
Vision-language models (VLMs) have made interactive digital museums increasingly feasible by connecting 3D digitization with natural-language artifact exploration. However, in cultural heritage domains such as ancient Greek pottery, reliable VLM assistance is limited by two challenges. First, open-ended interpretation requires grounding fine-grained 2D/3D visual evidence in specialized curatorial knowledge, yet the retrieval process may introduce weak sources and unverifiable references. Second, when the available evidence is incomplete, noisy, or ambiguous, VLMs often produce confident but unsupported answers instead of calibrated uncertainty. To address these challenges, we propose VaseMuseum, a lightweight and modular multimodal agent framework for intelligent digital museums of ancient Greek pottery. VaseMuseum combines an interactive virtual museum with VaseAgent, which supports both 2D images and 3D artifacts through multimodal perception, 3D-aware reasoning, external knowledge retrieval, and inference-time reliability control. Specifically, VaseAgent retrieves evidence from authoritative web and museum knowledge sources, and source-level control selects diverse and verifiable evidence before generation. Meanwhile, response-level control checks generated claims against the evidence pool and encourages neutral, evidence-bounded answers when support is insufficient or conflicting. Moreover, a training-free GRPO-style selection mechanism favors responses with valid references and calibrated confidence without updating the VLM backbone. Experiments in a realistic digital museum simulation show that VaseMuseum improves citation validity, reduces hallucinations on knowledge-intensive queries, and produces more neutral answers under ambiguity compared with search-enabled VLM baselines.