ChatPaper.aiChatPaper

StyleForge:超圖場域中的反事實推理實現室內家具風格化

StyleForge: Indoor Furniture Styling by Counterfactual Reasoning in a Hypergraph Field

August 3, 2026
作者: Lingwei Dang, Shishuo Shang, Pan Liu, Jiajia Cheng, Ziyan Qiu, Zhenhao Zhang, Yufei Zhu, Shenghui Huang, Qingxin Xiao, Yun Hao, Juntong Li, Qingyao Wu
cs.AI

摘要

固定佈局室內家具風格化需要在保持指定家具類別、位置、方向或尺度不變的前提下,選取能構成協調房間的資產。現有方法通常獨立檢索每個資產,或依賴靜態的局部關係,因此在場景合成後容易產生形狀、材質與色彩方面的衝突。我們提出 StyleForge,一種建構於動態超圖風格場域之上的場景級結構化選取框架。一個凍結的多模態大型語言模型從開放式風格請求與固定佈局中抽取結構化風格先驗,同時 StyleForge 為每個家具槽位維護一個可學習的候選分佈。在目標風格的條件下,動態超圖風格場域會自適應地啟動並加權由佈局引發的超邊,以捕捉家具之間的高階依賴關係。反事實風格偏好學習隨後將每個候選視為當前風格場域中的局部替換,並使用馬氏能量評估其情境相容性。訓練過程在最佳化風格場域與最佳化候選 logits 之間交替進行。在推論時,模型保持凍結,測試時訓練僅更新特定於房間的候選 logits,並隨著整體場景脈絡的演進逐步修正跨槽位的風格衝突。在 3D-FRONT 上的實驗展示了最先進的家具檢索與場景級風格一致性,相較於物件級與場景級檢索基準,能產生更為一致的固定佈局家具配置。
English
Fixed-layout indoor furniture styling requires selecting assets that form a coherent room without changing the prescribed furniture categories, positions, orientations, or scales. Existing approaches typically retrieve each asset independently or rely on static local relations, making them prone to shape, material, and color conflicts after scene composition. We introduce StyleForge, a scene-level structured selection framework built on a dynamic hypergraph style field. A frozen multimodal large language model extracts structured style priors from an open-ended style request and the fixed layout, while StyleForge maintains a learnable candidate distribution for each furniture slot. Conditioned on the target style, the dynamic hypergraph style field adaptively activates and weights layout-induced hyperedges to capture higher-order dependencies among furniture. Counterfactual style preference learning then treats each candidate as a local substitution in the current style field and evaluates its contextual compatibility using Mahalanobis energies. Training alternates between optimizing the style field and the candidate logits. At inference, the model remains frozen and test-time training updates only room-specific candidate logits, progressively correcting cross-slot style conflicts as the global scene context evolves. Experiments on 3D-FRONT demonstrate state-of-the-art furniture retrieval and scene-level style coherence, producing more coherent fixed-layout furniture arrangements than object- and scene-level retrieval baselines.