ChatPaper.aiChatPaper

MentalThink: メンタルSVG世界における思考の形成

MentalThink: Shaping Thoughts in Mental SVG World

July 3, 2026
著者: Kangheng Lin, Jisheng Yin, Dingming Li, En Yu, Yana Wei, Han Zhou, Liang Zhao, Hongyu Zhou, Hongbo Peng, Jianjian Sun, Zheng Ge, Xiangyu Zhang, Daxin Jiang, Jingyu Wang
cs.AI

要旨

本稿では、MentalThinkを提案する。これは、マルチモーダル大規模言語モデル(MLLM)に「心的」可視化のための実行可能なメカニズムを備える、視覚-記号推論パラダイムである。MentalThinkの中核はthink-with-SVGパイプラインであり、モデルはスケーラブルベクターグラフィックス(SVG)コードを中間視覚表現として生成、レンダリング、解釈することを学習し、複数ターンにわたる推論を行う。構造化されたベクタースケッチを作成することで、モデルは空間仮説を外在化し、決定的レンダリングを通じてそれを検証し、制約された幾何空間内で推論することが可能となり、人間の心的イメージプロセスを効果的に模倣する。本パラダイムは2段階の学習フレームワークによって具体化される。すなわち、SVGの構文整合性を担う教師ありファインチューニング(SFT)と、中間視覚仮説の反復的な検査、修正、洗練を促す多ターン強化学習(RL)を組み合わせる。広範な評価により、MentalThinkは空間理解および推論ベンチマーク(例:VSIBenchで55.1%、MindCubeで76.0%)において優れた性能を達成し、実行可能なベクターグラフィックスが動的視点取得、視覚的省察、および構成シーンの構築のための検証可能な視覚的ワークスペースを提供することが示された。
English
We introduce MentalThink, a visual-symbolic reasoning paradigm that equips Multimodal LLMs (MLLMs) with an executable mechanism for "mental" visualization. The core of MentalThink is a think-with-SVG pipeline, where the model learns to generate, render, and interpret scalable vector graphics (SVG) code as an intermediate visual representation for multi-turn reasoning. By creating structured vector sketches, the model can externalize spatial hypotheses, inspect them through deterministic rendering, and reason within a constrained geometric space, effectively mimicking the human process of mental imagery. We instantiate this paradigm through a two-stage training framework, combining Supervised Fine-Tuning (SFT) for SVG syntactic alignment with multi-turn Reinforcement Learning (RL) to encourage iterative inspection, revision, and refinement of intermediate visual hypotheses. Extensive evaluations demonstrate that MentalThink achieves superior performance on spatial understanding and reasoning benchmarks (e.g., 55.1% on VSIBench, 76.0% on MindCube), showing that executable vector graphics provide a verifiable visual workspace for dynamic perspective taking, visual reflection, and compositional scene construction.