ChatPaper.aiChatPaper

MentalThink : Façonner les pensées dans le monde SVG mental

MentalThink: Shaping Thoughts in Mental SVG World

July 3, 2026
Auteurs: Kangheng Lin, Jisheng Yin, Dingming Li, En Yu, Yana Wei, Han Zhou, Liang Zhao, Hongyu Zhou, Hongbo Peng, Jianjian Sun, Zheng Ge, Xiangyu Zhang, Daxin Jiang, Jingyu Wang
cs.AI

Résumé

Nous présentons MentalThink, un paradigme de raisonnement visuo-symbolique qui dote les LLM multimodaux (MLLM) d'un mécanisme exécutable de visualisation « mentale ». Le cœur de MentalThink est un pipeline de type « penser avec SVG », où le modèle apprend à générer, rendre et interpréter du code de graphiques vectoriels évolutifs (SVG) en tant que représentation visuelle intermédiaire pour un raisonnement à plusieurs tours. En créant des esquisses vectorielles structurées, le modèle peut externaliser des hypothèses spatiales, les inspecter via un rendu déterministe et raisonner dans un espace géométrique contraint, imitant ainsi efficacement le processus humain d'imagerie mentale. Nous concrétisons ce paradigme par un cadre d'apprentissage en deux étapes, combinant un ajustement fin supervisé (SFT) pour l'alignement syntaxique SVG avec un apprentissage par renforcement (RL) multi-tours afin d'encourager l'inspection, la révision et l'affinement itératifs des hypothèses visuelles intermédiaires. Des évaluations approfondies montrent que MentalThink atteint des performances supérieures sur des benchmarks de compréhension et de raisonnement spatial (par exemple, 55,1 % sur VSIBench, 76,0 % sur MindCube), démontrant que les graphiques vectoriels exécutables offrent un espace de travail visuel vérifiable pour la prise de perspective dynamique, la réflexion visuelle et la construction compositionnelle de scènes.
English
We introduce MentalThink, a visual-symbolic reasoning paradigm that equips Multimodal LLMs (MLLMs) with an executable mechanism for "mental" visualization. The core of MentalThink is a think-with-SVG pipeline, where the model learns to generate, render, and interpret scalable vector graphics (SVG) code as an intermediate visual representation for multi-turn reasoning. By creating structured vector sketches, the model can externalize spatial hypotheses, inspect them through deterministic rendering, and reason within a constrained geometric space, effectively mimicking the human process of mental imagery. We instantiate this paradigm through a two-stage training framework, combining Supervised Fine-Tuning (SFT) for SVG syntactic alignment with multi-turn Reinforcement Learning (RL) to encourage iterative inspection, revision, and refinement of intermediate visual hypotheses. Extensive evaluations demonstrate that MentalThink achieves superior performance on spatial understanding and reasoning benchmarks (e.g., 55.1% on VSIBench, 76.0% on MindCube), showing that executable vector graphics provide a verifiable visual workspace for dynamic perspective taking, visual reflection, and compositional scene construction.