PolicyShiftGuard: Benchmarking en Verbetering van Beleidsadaptieve Beeldveiligingsrichtlijnen
PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails
July 7, 2026
Auteurs: Mingyang Song, Luxin Xu, Haoyu Sun, Minzhou Pan, Yu Cheng, Bo Li
cs.AI
Samenvatting
Beeldbeveiligingssystemen (guardrails) worden doorgaans getraind en geëvalueerd onder een vast veiligheidsbeleid, waarbij veiligheid impliciet wordt behandeld als een intrinsieke eigenschap van een afbeelding. In de praktijk is dit anders: dezelfde afbeelding kan in het ene product worden toegestaan, in een ander worden beperkt, en opnieuw worden geweigerd wanneer een beleidsgrens verandert. We bestuderen beleidsadaptieve beeldbeveiliging, waarbij een model moet beslissen of een afbeelding het op dat moment geldende beleid schendt en moet generaliseren naar niet eerder geziene beleidsdefinities. We introduceren PolicyShiftBench, een uitgebreide benchmark met 2.000 beleidsonderscheidende instanties over 265 afbeeldingen, waarbij elke afbeelding gemiddeld wordt gecombineerd met 7,55 beleidsgeconditioneerde prompts om te testen of modellen zich aanpassen aan het actieve beleid in plaats van te vertrouwen op veiligheidsprioriteiten op afbeeldingsniveau. Vervolgens stellen we PolicyShiftGuard voor, een compacte beleidsgeconditioneerde guardrail getraind met een tweefasige trainingsmethode die Randomized Policy SFT (RP-SFT) combineert met Boundary-Pair Policy Adaptation (BP-Adapt). BP-Adapt traint overeenkomende prompts voor dezelfde afbeelding en risicocategorie met behulp van standaard labelsupervisie en een paarsgewijs vergelijkingsverlies dat blokkerende beleidsregels scheidt van goedkeurende beleidsregels. Experimenten tonen aan dat bestaande VLMs en gespecialiseerde guardrails kwetsbaar blijven bij beleidsverschuivingen, terwijl PolicyShiftGuard de beleidsgevoelige prestaties aanzienlijk verbetert. Het 7B-model behaalt een state-of-the-art prestatie van 76,9 gemiddelde F1 en 72,1 gemiddelde PSS op PolicyShiftBench, presteert goed op UnSafeBench en SafeEditBench, en verbetert de afweging tussen latentie en prestaties met een beknopt uitvoerformaat. Ablatiestudies bevestigen dat overeenkomende goedkeurings-/blokkeringsgrens paren essentieel zijn voor stabiele beleidsaanpassing.
English
Image guardrails are typically trained and evaluated under a fixed safety policy, implicitly treating safety as an intrinsic property of an image. Real deployments are different: the same image may be allowed in one product, restricted in another, and newly disallowed when a policy boundary changes. We study policy-adaptive image guardrailing, where a model must decide whether an image violates the currently supplied policy and generalize to held-out policy definitions. We introduce PolicyShiftBench, a comprehensive benchmark with 2,000 policy-discriminative instances over 265 images, where each image is paired with 7.55 policy-conditioned prompts on average to test whether models adapt to the active policy rather than relying on image-level safety priors. We then propose PolicyShiftGuard, a compact policy-conditioned guardrail trained with a two-stage training recipe that combines Randomized Policy SFT (RP-SFT) with Boundary-Pair Policy Adaptation (BP-Adapt). BP-Adapt trains matched prompts for the same image and risk category using standard label supervision and a pairwise comparison loss that separates blocking policies from passing policies. Experiments show that existing VLMs and specialized guardrails remain brittle under policy shifts, while PolicyShiftGuard substantially improves policy-sensitive performance. The 7B model achieves SOTA performance of 76.9 Avg. F1 and 72.1 Avg. PSS on PolicyShiftBench, transfers well to UnSafeBench and SafeEditBench, and improves the latency-performance trade-off with a concise output format. Ablations confirm that matched pass/block boundary pairs are essential for stable policy adaptation.