Guidage par Chaîne de Pensée Agentique pour un Raisonnement LLM Efficace et Contrôlable
Agentic Chain-of-Thought Steering for Efficient and Controllable LLM Reasoning
June 2, 2026
Auteurs: Yu Xia, Zhouhang Xie, Xin Xu, Byungkyu Kang, Prarit Lamba, Xiang Gao, Julian McAuley
cs.AI
Résumé
Les grands modèles de langage améliorent la précision des réponses finales grâce à un raisonnement étendu en chaîne de pensée, mais ils utilisent souvent les jetons de manière inefficace et offrent peu de contrôle en temps d'inférence. Les méthodes existantes de raisonnement efficace contrôlent la longueur de réflexion en raccourcissant, en arrêtant prématurément ou en compressant les traces, laissant implicite la manière dont le modèle raisonne. Dans cet article, nous proposons le pilotage agentique de la chaîne de pensée (ACTS), qui formule le pilotage du raisonnement comme un processus de décision markovien où un agent contrôleur adapte de manière adaptative un raisonneur figé pendant l'inférence. À chaque étape, le contrôleur observe la trace de raisonnement et le budget de réflexion restant, puis émet une action de pilotage consistant en une stratégie de raisonnement et une phrase de pilotage qui amorce l'étape suivante du raisonneur. Cela permet un contrôle stratégique tenant compte du budget pour un raisonnement efficace, tout en préservant la continuité de génération du raisonneur. Nous initialisons l'agent contrôleur à partir de nos trajectoires de pilotage synthétiques construites avec une augmentation multi-budget, et nous l'optimisons davantage via un apprentissage par renforcement avec une fonction de récompense conditionnée par le budget. Des expériences sur plusieurs bancs d'essai montrent qu'ACTS atteint des performances équivalentes à la réflexion complète avec des économies substantielles de jetons, et permet des compromis contrôlables entre précision et efficacité pour différents raisonneurs et tâches. Le code est disponible à l'adresse https://github.com/Andree-9/ACTS.
English
Large language models improve final-answer accuracy through extended chain-of-thought reasoning, but often spend tokens inefficiently and offer little inference-time control. Existing efficient reasoning methods control thinking length by shortening, early-stopping, or compressing traces, leaving how the model thinks implicit. In this paper, we propose Agentic Chain-of-Thought Steering (ACTS), which formulates reasoning steering as a Markov decision process where a controller agent adaptively steers a frozen reasoner during inference. At each step, the controller observes the reasoning trace and remaining thinking budget, then issues a steering action consisting of a reasoning strategy and a steering phrase that initiates the next reasoner step. This enables budget-aware strategy control for efficient reasoning while preserving the reasoner's generation continuity. We initialize the controller agent from our constructed synthetic steering trajectories with multi-budget augmentation, and further optimize it via reinforcement learning with budget-conditioned reward shaping. Experiments across multiple benchmarks show that ACTS matches full-thinking performance with substantial token savings, and enables controllable accuracy-efficiency trade-offs across different reasoners and tasks. The code is available at https://github.com/Andree-9/ACTS.