LEDGERMIND:具結構化證據帳本之來源約束多模態代理推理
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
July 30, 2026
作者: Enjun Du, Hange Zhou, Chenxu Du, Siyi Liu, Zirong Chen, Ziyu Zheng, Yongqi Zhang
cs.AI
摘要
用於視覺問答的多模態智能體日益以多步軌跡的形式運作,交錯執行感知、檢索與推理,然而評測方式仍大幅簡化為最終答案的準確率。此類聚合信號無法判別正確答案究竟是經由有依據的證據、語言先驗,抑或是偶然的錯誤相互抵消而達成。我們主張將多模態智能體軌跡視為受溯源約束的狀態機:工具輸出被標準化為結構化證據帳本,作為軌跡狀態;下游推理與決策主張僅能引用帳本中活躍的條目;接地檢查在實體與數值層級進行;修復則以型別化狀態轉換實現,此類轉換若缺乏工具產生的溯源資訊便無法引入新內容。我們將此設計具體化為 LedgerMind(搭載結構化證據帳本、受溯源約束的多模態智能體推理),並輔以三層接地協議、根據問題複雜度匹配推理深度的自適應雙路徑調度器,以及具備形式化溯源非放大保證的事件觸發驗證與修復引擎。我們運用 LedgerMind 針對四種反覆出現、且常被最終答案準確率掩蓋的失敗模式:缺乏支持的中間推理、具引用背書的實體幻覺(Phantom Grounding)、對簡單查詢的過度推理,以及修復時的資訊放大。跨多個多模態推理基準與骨幹多模態大型語言模型的實驗顯示,LedgerMind 同時提升了答案準確率與軌跡層級的忠實度。
English
Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate signal cannot tell whether a correct answer was reached through grounded evidence, language priors, or accidental error cancellation. We propose to treat a multimodal agent trajectory as a provenance-constrained state machine: tool outputs are normalized into a Structured Evidence Ledger that serves as the trajectory state, downstream reasoning and decision claims may cite only active ledger entries, grounding is checked at the entity and numeric level, and repair is realized as typed state transitions that cannot introduce content without tool-produced provenance. We instantiate this design as LedgerMind (Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger), augmented by a Three-Layer Grounding Protocol, an Adaptive Dual-Path Dispatcher that matches reasoning depth to question complexity, and an Event-Triggered Verification-and-Repair engine with a formal provenance non-amplification guarantee. We use LedgerMind to target four recurring failure patterns that final-answer accuracy tends to obscure: unsupported intermediate reasoning, citation-backed entity hallucination (Phantom Grounding), over-reasoning on simple queries, and repair-time amplification. Experiments across multiple multimodal reasoning benchmarks and backbone MLLMs show that LedgerMind improves both answer accuracy and trajectory-level faithfulness.