ChatPaper.aiChatPaper

정확성을 넘어서: 하이브리드 사고 기반 MLLM의 응답 행동 벤치마킹 및 정렬

Beyond Correctness: Benchmarking and Aligning Response Behaviors in Hybrid-Thinking MLLMs

August 17, 2026
저자: Xinming Wang, Weinong Wang, Hongming Yang, Yansong Lin, Zheng Ruan, Shangpin Peng, Qiming Peng, Nan Qiao, Fengyuan Lu, Guoqing Ma, Marito Li, Songyang Zhang, Saiyong Yang, Han Hu, Yonglong Tian, Xu-Yao Zhang
cs.AI

초록

하이브리드 사고 멀티모달 대규모 언어 모델(MLLM)은 단일 모델이 숙고적 사고와 지연 시간 효율적인 비사고 추론 간을 전환할 수 있게 한다. 이러한 모드들은 추론 예산에서 차이가 있지만, 산출되는 응답은 동일한 사용자 대면 기준을 충족해야 한다. 정확성만으로는 이러한 응답 품질을 특성화하기에 충분하지 않을 수 있다. 따라서 본 연구에서는 과업 정확도와 응답 패턴 오류를 상호 보완적 결과로 평가한다. 본 연구는 응답 패턴 정렬, 즉 사고 방식과 비사고 방식 인터페이스가 수용 가능한 최종 응답 행동을 유지하는지 여부를 통해 이러한 격차를 분석한다. 본 연구는 시각적 지각 및 그라운딩, 구조화된 이미지 이해, 멀티모달 지식 추론을 포괄하는 2,415개의 멀티모달 프롬프트로 구성된 오류 강화 진단 벤치마크인 PatternEval을 소개한다. PatternEval은 사고 사슬 누출, 응답 반복, 논리적 모순, 수행적 추론이라는 네 가지 반복적 오류를 검증한다. 응답 패턴 오류는 서로 다른 제공업체의 모델 전반에 걸쳐 광범위하게 나타나며, 비사고 추론은 상당히 더 높은 오류율을 보여 사고 방식과 비사고 방식 인터페이스 간의 체계적 정렬 불일치를 초래한다. 이러한 진단에 기반하여 본 연구는 응답 수준 보상 모델인 PatternRM과 강화 학습 중 패턴 특화 페널티를 도입하는 PatternRL을 개발한다. Qwen3-VL-4B 및 Qwen3-VL-8B에 대한 실험은 강화 학습에 패턴 특화 페널티를 통합하면 경미한 과업 성능 상충을 감수하면서 모드 간 정렬 불일치를 완화할 수 있음을 보여준다. PatternEval과 PatternRL은 함께 하이브리드 사고 인터페이스 전반에서 사용자에게 드러나는 응답 패턴을 정렬하기 위한 평가 및 훈련 프레임워크를 제공한다.
English
Hybrid-thinking multimodal large language models (MLLMs) allow a single model to alternate between deliberative thinking and latency-efficient non-thinking inference. Although these modes differ in reasoning budget, their delivered responses should satisfy the same user-facing standard. Correctness alone may not characterize this response quality; we therefore evaluate task accuracy and response-pattern failures as complementary outcomes. We study this gap through response-pattern alignment: whether thinking and non-thinking interfaces preserve acceptable final-response behavior. We introduce PatternEval, a failure-enriched diagnostic benchmark comprising 2,415 multimodal prompts spanning visual perception and grounding, structured image understanding, and multimodal knowledge reasoning. PatternEval tests four recurrent failures: chain-of-thought leakage, response repetition, logical contradiction, and performative reasoning. Response-pattern failures are widespread across models from different providers, with non-thinking inference exhibiting substantially higher failure rates and thereby creating systematic misalignment between thinking and non-thinking interfaces. Motivated by this diagnosis, we develop PatternRM, a response-level reward model, and PatternRL, which introduces pattern-specific penalties during reinforcement learning. Experiments on Qwen3-VL-4B and Qwen3-VL-8B show that incorporating pattern-specific penalties into reinforcement learning can mitigate cross-mode misalignment while incurring a marginal task performance trade-off. Together, PatternEval and PatternRL provide an evaluation-and-training framework for aligning user-visible response patterns across hybrid-thinking interfaces.