ChatPaper.aiChatPaper

VaseMuseum: 古代ギリシャ陶器のデジタルインテリジェント博物館

VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery

July 7, 2026
著者: Jiazi Wang, Nonghai Zhang, Qiushi Xie, Zeyu Zhang, Yufeng Chen, Yang Zhao, Ling Shao, Hao Tang
cs.AI

要旨

視覚言語モデル(VLM)は、3Dデジタル化と自然言語による遺物探索を結び付けることで、対話型デジタルミュージアムの実現性を高めてきた。しかし、古代ギリシャ陶器などの文化遺産分野では、信頼性の高いVLM支援は2つの課題によって制限されている。第一に、自由形式の解釈には、細粒度の2D/3D視覚的証拠を専門的な学芸知識に根拠付ける必要があるが、検索プロセスによって弱い情報源や検証不可能な参照が導入される可能性がある。第二に、利用可能な証拠が不完全、ノイズが多い、または曖昧な場合、VLMは調整された不確実性を示す代わりに、確信を持っているが根拠のない回答を生成することが多い。これらの課題に対処するため、我々はVaseMuseumを提案する。これは、古代ギリシャ陶器のインテリジェントデジタルミュージアムのための軽量かつモジュール型のマルチモーダルエージェントフレームワークである。VaseMuseumは、インタラクティブな仮想ミュージアムとVaseAgentを組み合わせる。VaseAgentは、マルチモーダル認識、3D認識推論、外部知識検索、および推論時の信頼性制御を通じて、2D画像と3D遺物の両方をサポートする。具体的には、VaseAgentは権威あるWebおよび博物館知識ソースから証拠を検索し、情報源レベルの制御によって、生成前に多様で検証可能な証拠を選択する。同時に、応答レベルの制御は、生成された主張を証拠プールに対して検証し、証拠が不十分または矛盾する場合には、中立的でエビデンスに基づいた回答を促す。さらに、学習不要のGRPOスタイル選択メカニズムは、VLMバックボーンを更新することなく、有効な参照と較正された信頼度を持つ応答を優先する。現実的なデジタルミュージアムシミュレーションでの実験により、VaseMuseumは、検索対応VLMベースラインと比較して、引用の妥当性を改善し、知識集約的なクエリにおけるハルシネーションを低減し、曖昧さの下でより中立的な回答を生成することが示された。
English
Vision-language models (VLMs) have made interactive digital museums increasingly feasible by connecting 3D digitization with natural-language artifact exploration. However, in cultural heritage domains such as ancient Greek pottery, reliable VLM assistance is limited by two challenges. First, open-ended interpretation requires grounding fine-grained 2D/3D visual evidence in specialized curatorial knowledge, yet the retrieval process may introduce weak sources and unverifiable references. Second, when the available evidence is incomplete, noisy, or ambiguous, VLMs often produce confident but unsupported answers instead of calibrated uncertainty. To address these challenges, we propose VaseMuseum, a lightweight and modular multimodal agent framework for intelligent digital museums of ancient Greek pottery. VaseMuseum combines an interactive virtual museum with VaseAgent, which supports both 2D images and 3D artifacts through multimodal perception, 3D-aware reasoning, external knowledge retrieval, and inference-time reliability control. Specifically, VaseAgent retrieves evidence from authoritative web and museum knowledge sources, and source-level control selects diverse and verifiable evidence before generation. Meanwhile, response-level control checks generated claims against the evidence pool and encourages neutral, evidence-bounded answers when support is insufficient or conflicting. Moreover, a training-free GRPO-style selection mechanism favors responses with valid references and calibrated confidence without updating the VLM backbone. Experiments in a realistic digital museum simulation show that VaseMuseum improves citation validity, reduces hallucinations on knowledge-intensive queries, and produces more neutral answers under ambiguity compared with search-enabled VLM baselines.