VaseMuseum: 고대 그리스 도자기를 위한 디지털 지능형 박물관
VaseMuseum: Digital Intelligent Museum for Ancient Greek Pottery
July 7, 2026
저자: Jiazi Wang, Nonghai Zhang, Qiushi Xie, Zeyu Zhang, Yufeng Chen, Yang Zhao, Ling Shao, Hao Tang
cs.AI
초록
시각-언어 모델(VLM)은 3D 디지털화와 자연어 기반 유물 탐색을 연결함으로써 대화형 디지털 박물관을 점점 더 실현 가능하게 만들고 있다. 그러나 고대 그리스 도자기와 같은 문화유산 분야에서 신뢰할 수 있는 VLM 지원은 두 가지 과제로 인해 제한된다. 첫째, 개방형 해석은 세밀한 2D/3D 시각 증거를 전문 큐레이터 지식에 근거하여 정립해야 하지만, 검색 과정에서 약한 출처와 검증 불가능한 참조가 도입될 수 있다. 둘째, 이용 가능한 증거가 불완전하거나 잡음이 많거나 모호할 때, VLM은 보정된 불확실성 대신 자신감은 있지만 뒷받침되지 않는 답변을 생성하는 경우가 많다. 이러한 과제를 해결하기 위해 우리는 고대 그리스 도자기를 위한 지능형 디지털 박물관용 경량 모듈형 멀티모달 에이전트 프레임워크인 VaseMuseum을 제안한다. VaseMuseum은 대화형 가상 박물관과 VaseAgent를 결합하며, VaseAgent는 멀티모달 인식, 3D 인식 추론, 외부 지식 검색, 추론 시 신뢰도 제어를 통해 2D 이미지와 3D 유물을 모두 지원한다. 구체적으로 VaseAgent는 권위 있는 웹 및 박물관 지식 소스에서 증거를 검색하며, 생성 전에 소스 수준 제어를 통해 다양하고 검증 가능한 증거를 선택한다. 한편, 응답 수준 제어는 생성된 주장을 증거 풀과 대조하여 확인하고, 증거가 불충분하거나 상충될 때 중립적이고 증거에 기반한 답변을 유도한다. 또한, 훈련 없는 GRPO 스타일 선택 메커니즘은 VLM 백본을 업데이트하지 않고도 유효한 참조와 보정된 신뢰도를 가진 응답을 선호한다. 실제적인 디지털 박물관 시뮬레이션에서의 실험 결과, VaseMuseum은 검색이 가능한 VLM 베이스라인에 비해 인용 타당성을 개선하고, 지식 집약적 질의에서 환각(허구 생성)을 줄이며, 모호한 상황에서 더 중립적인 답변을 생성함을 보여준다.
English
Vision-language models (VLMs) have made interactive digital museums increasingly feasible by connecting 3D digitization with natural-language artifact exploration. However, in cultural heritage domains such as ancient Greek pottery, reliable VLM assistance is limited by two challenges. First, open-ended interpretation requires grounding fine-grained 2D/3D visual evidence in specialized curatorial knowledge, yet the retrieval process may introduce weak sources and unverifiable references. Second, when the available evidence is incomplete, noisy, or ambiguous, VLMs often produce confident but unsupported answers instead of calibrated uncertainty. To address these challenges, we propose VaseMuseum, a lightweight and modular multimodal agent framework for intelligent digital museums of ancient Greek pottery. VaseMuseum combines an interactive virtual museum with VaseAgent, which supports both 2D images and 3D artifacts through multimodal perception, 3D-aware reasoning, external knowledge retrieval, and inference-time reliability control. Specifically, VaseAgent retrieves evidence from authoritative web and museum knowledge sources, and source-level control selects diverse and verifiable evidence before generation. Meanwhile, response-level control checks generated claims against the evidence pool and encourages neutral, evidence-bounded answers when support is insufficient or conflicting. Moreover, a training-free GRPO-style selection mechanism favors responses with valid references and calibrated confidence without updating the VLM backbone. Experiments in a realistic digital museum simulation show that VaseMuseum improves citation validity, reduces hallucinations on knowledge-intensive queries, and produces more neutral answers under ambiguity compared with search-enabled VLM baselines.