ChatPaper.aiChatPaper

LEDGERMIND:構造化エビデンス台帳を用いた来歴制約付きマルチモーダル・エージェンティック推論

LEDGERMIND: Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger

July 30, 2026
著者: Enjun Du, Hange Zhou, Chenxu Du, Siyi Liu, Zirong Chen, Ziyu Zheng, Yongqi Zhang
cs.AI

要旨

視覚的質問応答のためのマルチモーダルエージェントは、知覚・検索・推論を織り交ぜたマルチステップ軌跡として動作することが増えているが、評価は依然として大部分が最終回答の正解率に還元されている。この集約的な指標は、正解が裏付けられたエビデンス、言語的な事前知識、あるいは偶然の誤差の相殺によって達成されたのかを判別できない。我々は、マルチモーダルエージェント軌跡を来歴制約付き状態機械として扱うことを提案する。すなわち、ツール出力は軌跡状態として機能する構造化エビデンス台帳(Structured Evidence Ledger)に正規化され、後続の推論および決定の主張はアクティブな台帳エントリのみを引用でき、グラウンディングはエンティティおよび数値レベルで検証され、修復は型付き状態遷移として実現され、ツール生成の来歴なしに内容を導入することはできない。我々はこの設計を、LedgerMind(構造化エビデンス台帳を用いた来歴制約付きマルチモーダルエージェント推論)として具体化する。これは、三層グラウンディングプロトコル(Three-Layer Grounding Protocol)、質問の複雑さに応じて推論の深さを調整する適応的二経路ディスパッチャ(Adaptive Dual-Path Dispatcher)、および形式的な来歴非増幅保証を備えたイベント駆動型検証・修復エンジン(Event-Triggered Verification-and-Repair engine)によって拡張される。我々はLedgerMindを用いて、最終回答の正解率では覆い隠されがちな四つの反復的失敗パターンに対処する:根拠のない中間推論、引用を伴うエンティティ幻覚(ファントム・グラウンディング)、単純なクエリへの過剰推論、および修復時の増幅である。複数のマルチモーダル推論ベンチマークとバックボーンMLLM(マルチモーダル大規模言語モデル)にわたる実験により、LedgerMindが回答精度と軌跡レベルの忠実性の両方を改善することが示される。
English
Multimodal agents for visual question answering increasingly operate as multi-step trajectories that interleave perception, retrieval, and reasoning, yet evaluation still largely reduces to final-answer accuracy. This aggregate signal cannot tell whether a correct answer was reached through grounded evidence, language priors, or accidental error cancellation. We propose to treat a multimodal agent trajectory as a provenance-constrained state machine: tool outputs are normalized into a Structured Evidence Ledger that serves as the trajectory state, downstream reasoning and decision claims may cite only active ledger entries, grounding is checked at the entity and numeric level, and repair is realized as typed state transitions that cannot introduce content without tool-produced provenance. We instantiate this design as LedgerMind (Provenance-Constrained Multimodal Agentic Reasoning with a Structured Evidence Ledger), augmented by a Three-Layer Grounding Protocol, an Adaptive Dual-Path Dispatcher that matches reasoning depth to question complexity, and an Event-Triggered Verification-and-Repair engine with a formal provenance non-amplification guarantee. We use LedgerMind to target four recurring failure patterns that final-answer accuracy tends to obscure: unsupported intermediate reasoning, citation-backed entity hallucination (Phantom Grounding), over-reasoning on simple queries, and repair-time amplification. Experiments across multiple multimodal reasoning benchmarks and backbone MLLMs show that LedgerMind improves both answer accuracy and trajectory-level faithfulness.