ChatPaper.aiChatPaper

正確性を超えて:ハイブリッド思考MLLMにおける応答挙動のベンチマーキングとアラインメント

Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs

August 17, 2026
著者: Xinming Wang, Weinong Wang, Hongming Yang, Yansong Lin, Zheng Ruan, Shangpin Peng, Qiming Peng, Nan Qiao, Fengyuan Lu, Guoqing Ma, Marito Li, Songyang Zhang, Saiyong Yang, Han Hu, Yonglong Tian, Xu-Yao Zhang
cs.AI

要旨

ハイブリッド思考マルチモーダル大規模言語モデル(MLLMs)は、単一のモデルが熟慮的思考と低遅延の非思考推論を切り替えることを可能にする。これらのモードは推論予算が異なるものの、生成される応答は同じユーザー向け基準を満たすべきである。応答品質は正しさのみでは特徴づけられない可能性があるため、我々はタスク精度と応答パターンの失敗を相補的な成果として評価する。我々はこのギャップを応答パターンの整合性を通じて研究する:すなわち、思考インターフェースと非思考インターフェースが許容可能な最終応答の振る舞いを維持するかどうかである。我々はPatternEvalを紹介する。これは、視覚知覚とグラウンディング、構造化画像理解、マルチモーダル知識推論にわたる2,415のマルチモーダルプロンプトから構成される、失敗事例を豊富に含む診断ベンチマークである。PatternEvalは、思考連鎖の漏出、応答の繰り返し、論理的矛盾、パフォーマティブ推論という4つの反復的失敗を検証する。応答パターンの失敗は異なる提供元のモデルに広く見られ、非思考推論は著しく高い失敗率を示し、これにより思考インターフェースと非思考インターフェースの間に系統的な不整合が生じる。この診断に動機づけられ、我々は応答レベルの報酬モデルであるPatternRMと、強化学習中にパターン固有のペナルティを導入するPatternRLを開発する。Qwen3-VL-4BとQwen3-VL-8Bを用いた実験は、パターン固有のペナルティを強化学習に組み込むことで、わずかなタスク性能のトレードオフを伴いつつ、モード間の不整合を軽減できることを示す。合わせて、PatternEvalとPatternRLは、ハイブリッド思考インターフェース間でユーザーに可視な応答パターンを整合させるための評価・訓練フレームワークを提供する。
English
Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative thinking and latency-efficient non-thinking inference. Although these modes differ in reasoning budget, their delivered responses should satisfy the same user-facing standard. Correctness alone may not characterize this response quality; we therefore evaluate task accuracy and response-pattern failures as complementary outcomes. We study this gap through response-pattern alignment: whether thinking and non-thinking interfaces preserve acceptable final-response behavior. We introduce PatternEval, a failure-enriched diagnostic benchmark comprising 2,415 multimodal prompts spanning visual perception and grounding, structured image understanding, and multimodal knowledge reasoning. PatternEval tests four recurrent failures: chain-of-thought leakage, response repetition, logical contradiction, and performative reasoning. Response-pattern failures are widespread across models from different providers, with non-thinking inference exhibiting substantially higher failure rates and thereby creating systematic misalignment between thinking and non-thinking interfaces. Motivated by this diagnosis, we develop PatternRM, a response-level reward model, and PatternRL, which introduces pattern-specific penalties during reinforcement learning. Experiments on Qwen3-VL-4B and Qwen3-VL-8B show that incorporating pattern-specific penalties into reinforcement learning can mitigate cross-mode misalignment while incurring a marginal task performance trade-off. Together, PatternEval and PatternRL provide an evaluation-and-training framework for aligning user-visible response patterns across hybrid-thinking interfaces.