ChatPaper.aiChatPaper

MentalThink:在思維的SVG世界中塑造想法

MentalThink: Shaping Thoughts in Mental SVG World

July 3, 2026
作者: Kangheng Lin, Jisheng Yin, Dingming Li, En Yu, Yana Wei, Han Zhou, Liang Zhao, Hongyu Zhou, Hongbo Peng, Jianjian Sun, Zheng Ge, Xiangyu Zhang, Daxin Jiang, Jingyu Wang
cs.AI

摘要

我們介紹 MentalThink——一種視覺-符號推理典範,賦予多模態大型語言模型(MLLMs)一套可執行的「心理」視覺化機制。MentalThink 的核心是一條「思考搭配 SVG」流程,讓模型學會生成、渲染並解釋可縮放向量圖形(SVG)程式碼,作為多輪推理的中間視覺表徵。透過建立結構化的向量草圖,模型能將空間假設外化,經由確定性渲染進行檢查,並在受限的幾何空間中推理,有效模仿人類心理意象的過程。我們透過兩階段訓練框架實例化此典範,結合監督式微調(SFT)來達成 SVG 語法對齊,以及多輪強化學習(RL)來鼓勵對中間視覺假設進行疊代檢查、修正與精煉。廣泛的評估結果顯示,MentalThink 在空間理解與推理基準(例如 VSIBench 達 55.1%、MindCube 達 76.0%)上表現卓越,證明了可執行向量圖形能為動態視角轉換、視覺反思與組合場景建構提供可驗證的視覺工作空間。
English
We introduce MentalThink, a visual-symbolic reasoning paradigm that equips Multimodal LLMs (MLLMs) with an executable mechanism for "mental" visualization. The core of MentalThink is a think-with-SVG pipeline, where the model learns to generate, render, and interpret scalable vector graphics (SVG) code as an intermediate visual representation for multi-turn reasoning. By creating structured vector sketches, the model can externalize spatial hypotheses, inspect them through deterministic rendering, and reason within a constrained geometric space, effectively mimicking the human process of mental imagery. We instantiate this paradigm through a two-stage training framework, combining Supervised Fine-Tuning (SFT) for SVG syntactic alignment with multi-turn Reinforcement Learning (RL) to encourage iterative inspection, revision, and refinement of intermediate visual hypotheses. Extensive evaluations demonstrate that MentalThink achieves superior performance on spatial understanding and reasoning benchmarks (e.g., 55.1% on VSIBench, 76.0% on MindCube), showing that executable vector graphics provide a verifiable visual workspace for dynamic perspective taking, visual reflection, and compositional scene construction.