ChatPaper.aiChatPaper

InstanceControl: Controleerbare complexe beeldgeneratie zonder instantielabeling

InstanceControl: Controllable Complex Image Generation without Instance Labeling

June 30, 2026
Auteurs: Xiaoyu Liu, Huan Wang, Fan Li, Zhixin Wang, Jiaqi Xu, Ming Liu, Wangmeng Zuo
cs.AI

Samenvatting

Controleerbare beeldgeneratiemethoden, zoals ControlNet, hebben een opmerkelijk vermogen getoond om visuele condities (bijv. dieptekaarten) te introduceren ter begeleiding van beeldgeneratie. Deze methoden hebben echter vaak moeite met complexe multi-instance scènes, wat leidt tot frequente attribuutverwarring tussen instanties. Hoewel recente benaderingen dit proberen te verhelpen via handmatige instantielabeling, is dergelijke vereiste arbeidsintensief. In dit artikel stellen wij InstanceControl voor, een nieuwe multi-instance controleerbare generatiemethode die de noodzaak van instantielabeling elimineert. Wij identificeren de voornaamste bottleneck in bestaande methoden als het onvermogen om instantiebeschrijvingen nauwkeurig te koppelen aan hun corresponderende gebieden binnen visuele condities. Om dit aan te pakken, maken wij gebruik van het Visie-Taal Model (VTM) om instantieniveau-correspondenties tot stand te brengen tussen tekstprompts en visuele condities. Specifiek parseert het VTM automatisch instantiebeschrijvingen uit de tekstprompts en voorspelt tegelijkertijd instantiemaskers op basis van de visuele condities. Bovendien, aangezien de voorspelde maskers ruis kunnen bevatten, introduceren wij een adaptieve maskerverfijningsstrategie die deze instantiemaskers dynamisch verfijnt tijdens het generatieproces. Uitgebreide experimenten tonen aan dat onze benadering beter presteert dan de nieuwste methoden, met superieure betrouwbaarheid en nauwkeurige instantieniveau-controle.
English
Controllable image generation methods, such as ControlNet, have demonstrated a remarkable capacity to introduce visual conditions(e.g., depth maps) to guide image generation. However, these methods often struggle with complex multi-instance scenes, frequently leading to attribute confusion among instances. While recent approaches attempt to mitigate this via manual instance labeling, such requirements are labor-intensive. In this paper, we propose InstanceControl, a novel multi-instance controllable generation method that eliminates the need for instance labeling. We identify the primary bottleneck in existing methods as the inability to accurately associate instance descriptions with their corresponding regions within visual conditions. To address this, we leverage the Vision-Language Model (VLM) to establish instance-level correspondences between text prompts and visual conditions. Specifically, the VLM automatically parses instance descriptions from the text prompts and simultaneously predicts instance masks based on the visual conditions. Furthermore, since the predicted masks may contain noise, we introduce an adaptive mask refinement strategy that dynamically refines these instance masks during the generation process. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods, achieving superior fidelity and precise instance-level control.