ChatPaper.aiChatPaper

StyleForge:超图场中基于反事实推理的室内家具风格化

StyleForge: Indoor Furniture Styling by Counterfactual Reasoning in a Hypergraph Field

August 3, 2026
作者: Lingwei Dang, Shishuo Shang, Pan Liu, Jiajia Cheng, Ziyan Qiu, Zhenhao Zhang, Yufei Zhu, Shenghui Huang, Qingxin Xiao, Yun Hao, Juntong Li, Qingyao Wu
cs.AI

摘要

固定布局室内家具造型需要在保持规定的家具类别、位置、朝向和尺度不变的前提下,选择能够形成协调一致房间的资产。现有方法通常独立检索每个资产或依赖静态局部关系,因此在场景合成后容易出现形状、材质和色彩冲突。我们提出 StyleForge,一种基于动态超图风格场的场景级结构化选择框架。一个冻结的多模态大语言模型从未限定的风格请求和固定布局中提取结构化风格先验,同时 StyleForge 为每个家具槽位维护一个可学习的候选分布。在目标风格条件下,动态超图风格场自适应地激活并加权由布局诱导的超边,以捕捉家具之间的高阶依赖关系。随后,反事实风格偏好学习将每个候选视为当前风格场中的局部替换,并使用马氏能量评估其上下文兼容性。训练过程交替优化风格场和候选逻辑值。在推理阶段,模型保持冻结,测试时训练仅更新房间特定的候选逻辑值,随着全局场景上下文不断演化,逐步修正跨槽位风格冲突。在 3D-FRONT 上的实验表明,该方法在家具检索和场景级风格一致性方面达到了最先进的水平,生成的固定布局家具布置比对象级和场景级检索基线更加协调一致。
English
Fixed-layout indoor furniture styling requires selecting assets that form a coherent room without changing the prescribed furniture categories, positions, orientations, or scales. Existing approaches typically retrieve each asset independently or rely on static local relations, making them prone to shape, material, and color conflicts after scene composition. We introduce StyleForge, a scene-level structured selection framework built on a dynamic hypergraph style field. A frozen multimodal large language model extracts structured style priors from an open-ended style request and the fixed layout, while StyleForge maintains a learnable candidate distribution for each furniture slot. Conditioned on the target style, the dynamic hypergraph style field adaptively activates and weights layout-induced hyperedges to capture higher-order dependencies among furniture. Counterfactual style preference learning then treats each candidate as a local substitution in the current style field and evaluates its contextual compatibility using Mahalanobis energies. Training alternates between optimizing the style field and the candidate logits. At inference, the model remains frozen and test-time training updates only room-specific candidate logits, progressively correcting cross-slot style conflicts as the global scene context evolves. Experiments on 3D-FRONT demonstrate state-of-the-art furniture retrieval and scene-level style coherence, producing more coherent fixed-layout furniture arrangements than object- and scene-level retrieval baselines.