ChatPaper.aiChatPaper

LEDGERMIND: 구조화된 증거 원장을 활용한 출처 제약적 멀티모달 에이전트 추론

LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger

July 30, 2026
저자: Enjun Du, Hange Zhou, Chenxu Du, Siyi Liu, Zirong Chen, Ziyu Zheng, Yongqi Zhang
cs.AI

초록

멀티모달 에이전트 기반 시각 질의 응답은 지각, 검색, 추론을 교차 수행하는 다단계 궤적으로 점점 더 작동하게 되었지만, 평가는 여전히 대부분 최종 답변 정확도로 축소된다. 이러한 종합 신호는 올바른 답변이 근거 기반 증거, 언어적 사전 지식, 또는 우연한 오류 상쇄를 통해 도달되었는지 구분할 수 없다. 우리는 멀티모달 에이전트 궤적을 출처 제약 상태 기계(provenance-constrained state machine)로 취급할 것을 제안한다. 도구 출력은 구조화된 증거 원장(Structured Evidence Ledger)으로 정규화되어 궤적 상태로 기능하며, 이후의 추론 및 결정 주장은 활성 원장 항목만 인용할 수 있고, 그라운딩은 개체 및 수치 수준에서 검사되며, 수리는 도구가 생성한 출처 없이는 내용을 도입할 수 없는 유형화된 상태 전이로 구현된다. 우리는 이 설계를 LedgerMind(구조화된 증거 원장을 사용하는 출처 제약 멀티모달 에이전트 추론)로 구현하고, 3계층 그라운딩 프로토콜, 질문 복잡도에 맞춰 추론 깊이를 조정하는 적응형 이중 경로 디스패처, 그리고 형식적 출처 비증폭 보장을 갖춘 이벤트 트리거 검증-수리 엔진으로 보강한다. LedgerMind는 최종 답변 정확도가 흔히 가려버리는 네 가지 반복적 실패 패턴, 즉 근거 없는 중간 추론, 인용 기반 개체 환각(팬텀 그라운딩), 단순 질문에 대한 과도한 추론, 수리 시점 증폭을 겨냥한다. 여러 멀티모달 추론 벤치마크와 백본 MLLM에 걸친 실험은 LedgerMind가 답변 정확도와 궤적 수준 충실도를 모두 향상시킴을 보여준다.
English
Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate signal cannot tell whether a correct answer was reached through grounded evidence, language priors, or accidental error cancellation. We propose to treat a multimodal agent trajectory as a provenance-constrained state machine: tool outputs are normalized into a Structured Evidence Ledger that serves as the trajectory state, downstream reasoning and decision claims may cite only active ledger entries, grounding is checked at the entity and numeric level, and repair is realized as typed state transitions that cannot introduce content without tool-produced provenance. We instantiate this design as LedgerMind (Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger), augmented by a Three-Layer Grounding Protocol, an Adaptive Dual-Path Dispatcher that matches reasoning depth to question complexity, and an Event-Triggered Verification-and-Repair engine with a formal provenance non-amplification guarantee. We use LedgerMind to target four recurring failure patterns that final-answer accuracy tends to obscure: unsupported intermediate reasoning, citation-backed entity hallucination (Phantom Grounding), over-reasoning on simple queries, and repair-time amplification. Experiments across multiple multimodal reasoning benchmarks and backbone MLLMs show that LedgerMind improves both answer accuracy and trajectory-level faithfulness.