VaseMuseum:古希腊陶器数字智能博物馆
VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery
July 7, 2026
作者: Jiazi Wang, Nonghai Zhang, Qiushi Xie, Zeyu Zhang, Yufeng Chen, Yang Zhao, Ling Shao, Hao Tang
cs.AI
摘要
视觉-语言模型(VLMs)通过将三维数字化与自然语言文物探索相结合,使交互式数字博物馆的可行性日益增强。然而,在诸如古希腊陶器等文化遗产领域,可靠的VLM辅助受限于两大挑战。首先,开放式解读要求将细粒度的二维/三维视觉证据锚定在专业的策展知识中,但检索过程可能引入弱来源和不可验证的引用。其次,当可用证据不完整、存在噪声或模糊不清时,VLMs往往给出自信但缺乏依据的答案,而非经过校准的不确定性。为应对这些挑战,我们提出VaseMuseum——一个面向古希腊陶器智能数字博物馆的轻量级、模块化多模态智能体框架。VaseMuseum将交互式虚拟博物馆与VaseAgent相结合,后者通过多模态感知、三维感知推理、外部知识检索以及推理时可靠性控制,同时支持二维图像和三维文物。具体而言,VaseAgent从权威网络和博物馆知识源中检索证据,并在生成前通过源级控制选择多样且可验证的证据。同时,响应级控制将生成的陈述与证据池进行核对,并在支持不足或存在矛盾时鼓励中立的、基于证据的答案。此外,一种免训练的GRPO风格选择机制更倾向于选择包含有效引用和校准置信度的响应,且无需更新VLM主干网络。在真实的数字博物馆模拟实验中,与支持检索的VLM基线相比,VaseMuseum提升了引用有效性,减少了知识密集型查询中的幻觉,并在模糊情境下生成更中立的答案。
English
Vision-language models (VLMs) have made interactive digital museums increasingly feasible by connecting 3D digitization with natural-language artifact exploration. However, in cultural heritage domains such as ancient Greek pottery, reliable VLM assistance is limited by two challenges. First, open-ended interpretation requires grounding fine-grained 2D/3D visual evidence in specialized curatorial knowledge, yet the retrieval process may introduce weak sources and unverifiable references. Second, when the available evidence is incomplete, noisy, or ambiguous, VLMs often produce confident but unsupported answers instead of calibrated uncertainty. To address these challenges, we propose VaseMuseum, a lightweight and modular multimodal agent framework for intelligent digital museums of ancient Greek pottery. VaseMuseum combines an interactive virtual museum with VaseAgent, which supports both 2D images and 3D artifacts through multimodal perception, 3D-aware reasoning, external knowledge retrieval, and inference-time reliability control. Specifically, VaseAgent retrieves evidence from authoritative web and museum knowledge sources, and source-level control selects diverse and verifiable evidence before generation. Meanwhile, response-level control checks generated claims against the evidence pool and encourages neutral, evidence-bounded answers when support is insufficient or conflicting. Moreover, a training-free GRPO-style selection mechanism favors responses with valid references and calibrated confidence without updating the VLM backbone. Experiments in a realistic digital museum simulation show that VaseMuseum improves citation validity, reduces hallucinations on knowledge-intensive queries, and produces more neutral answers under ambiguity compared with search-enabled VLM baselines.