LEDGERMIND:基于结构化证据账本的溯源约束多模态智能体推理
LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger
July 30, 2026
作者: Enjun Du, Hange Zhou, Chenxu Du, Siyi Liu, Zirong Chen, Ziyu Zheng, Yongqi Zhang
cs.AI
摘要
面向视觉问答的多模态智能体日益以多步轨迹形式运行,交织感知、检索与推理,然而评估仍主要归结为最终答案准确率。这一聚合信号无法判断正确答案究竟源于有依据的证据、语言先验,还是偶然的误差抵消。我们提出将多模态智能体轨迹视为一种溯源约束状态机:工具输出被规范化为结构化证据账本,作为轨迹状态;下游推理与决策断言仅可引用活动账本条目;锚定检查在实体与数值层面进行;修复以类型化状态转换的形式实现,若无工具生成的溯源则无法引入任何内容。我们将该设计实例化为LedgerMind(基于结构化证据账本的溯源约束多模态智能体推理),并辅以三层锚定协议、一个将推理深度与问题复杂度相匹配的自适应双路径调度器,以及一个具备形式化溯源非放大保证的事件触发式验证与修复引擎。我们利用LedgerMind针对最终答案准确率往往掩盖的四类反复出现的失败模式:无依据的中间推理、有引用支撑的实体幻觉(幻影锚定)、对简单查询的过度推理,以及修复时放大。在多个多模态推理基准和骨干多模态大语言模型上的实验表明,LedgerMind同时提升了答案准确率和轨迹级忠实性。
English
Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate signal cannot tell whether a correct answer was reached through grounded evidence, language priors, or accidental error cancellation. We propose to treat a multimodal agent trajectory as a provenance-constrained state machine: tool outputs are normalized into a Structured Evidence Ledger that serves as the trajectory state, downstream reasoning and decision claims may cite only active ledger entries, grounding is checked at the entity and numeric level, and repair is realized as typed state transitions that cannot introduce content without tool-produced provenance. We instantiate this design as LedgerMind (Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger), augmented by a Three-Layer Grounding Protocol, an Adaptive Dual-Path Dispatcher that matches reasoning depth to question complexity, and an Event-Triggered Verification-and-Repair engine with a formal provenance non-amplification guarantee. We use LedgerMind to target four recurring failure patterns that final-answer accuracy tends to obscure: unsupported intermediate reasoning, citation-backed entity hallucination (Phantom Grounding), over-reasoning on simple queries, and repair-time amplification. Experiments across multiple multimodal reasoning benchmarks and backbone MLLMs show that LedgerMind improves both answer accuracy and trajectory-level faithfulness.