ChatPaper.aiChatPaper

Prompt-Activatie Dualiteit: Verbetering van Activatiesturing via Interventies op Aandachtsniveau

Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions

May 11, 2026
Auteurs: Diancheng Kang, Zheyuan Liu, Ningshan Ma, Yue Huang, Zhaoxuan Tan, Meng Jiang
cs.AI

Samenvatting

Activeringssturing beheerst het gedrag van taalmodellen door tijdens de inferentie richtingen toe te voegen aan interne representaties, maar standaard residual-stroomsturing kan falen in toestandsafhankelijke dialoog. We identificeren KV-cacheverontreiniging als een belangrijke faalmodus: gestuurde token-toestanden worden opgeslagen en herhaaldelijk hergebruikt, waardoor een lokale perturbatie verandert in een cumulatieve coherentieverslechtering. Om deze uitdaging aan te pakken, stellen we Gated Cropped Attention-Delta-sturing (GCAD) voor, die stuurssignalen extraheert uit systeempromptbijdragen aan zelfaandacht en deze toepast met token-niveau-poortregeling. In experimenten met persoonlijkheidssturing behoudt GCAD de controle over eigenschappen terwijl het de coherentie op lange termijn aanzienlijk verbetert. Op de belangrijkste multi-turn-benchmark verbetert GCAD de gemiddelde coherentie-afwijking van -18,6 naar -1,9 en verhoogt het de trekexpressie op beurt 10 van 78,0 naar 93,1. Deze resultaten suggereren dat activeringssturing betrouwbaarder wordt wanneer interventies de prompt-gemedieerde paden volgen die modellen al gebruiken voor gedragscontrole.
English
Activation steering controls language model behavior by adding directions to internal representations at inference time, but standard residual-stream steering can fail in stateful dialogue. We identify KV-cache contamination as a key failure mode: steered token states are stored and repeatedly reused, turning a local perturbation into cumulative coherence degradation. To address this challenge, we propose Gated Cropped Attention-Delta steering (GCAD), which extracts steering signals from system-prompt contributions to self-attention and applies them with token-level gating. Across persona-steering experiments, GCAD preserves trait control while substantially improving long-horizon coherence. On the main multi-turn benchmark, GCAD improves average coherence drift from -18.6 to -1.9 and raises turn-10 trait expression from 78.0 to 93.1. These results suggest that activation steering becomes more reliable when interventions follow the prompt-mediated pathways that models already use for behavioral control.