StyleForge: ハイパーグラフ場における反事実的推論による屋内家具スタイリング
StyleForge: Indoor Furniture Styling by Counterfactual Reasoning in a Hypergraph Field
August 3, 2026
著者: Lingwei Dang, Shishuo Shang, Pan Liu, Jiajia Cheng, Ziyan Qiu, Zhenhao Zhang, Yufei Zhu, Shenghui Huang, Qingxin Xiao, Yun Hao, Juntong Li, Qingyao Wu
cs.AI
要旨
固定レイアウトでの屋内家具スタイリングでは、指定された家具カテゴリ、位置、向き、スケールを変更せずに、一貫性のある部屋を構成するアセットを選択することが求められる。既存手法は通常、各アセットを独立に検索するか、静的な局所関係に依存するため、シーン合成後に形状・素材・色の矛盾が生じやすい。本稿では、動的ハイパーグラフスタイルフィールドに基づくシーンレベルの構造化選択フレームワークであるStyleForgeを提案する。凍結されたマルチモーダル大規模言語モデルが、自由形式のスタイル要求と固定レイアウトから構造化されたスタイル事前知識を抽出し、StyleForgeは各家具スロットに対して学習可能な候補分布を維持する。対象スタイルを条件として、動的ハイパーグラフスタイルフィールドはレイアウト由来のハイパーエッジを適応的に活性化・重み付けし、家具間の高次依存関係を捉える。反事実的スタイル嗜好学習では、各候補を現在のスタイルフィールドにおける局所的置換として扱い、マハラノビスエネルギーを用いて文脈的互換性を評価する。学習はスタイルフィールドと候補ロジットの最適化を交互に行う。推論時には、モデルは凍結されたままであり、テスト時学習は部屋固有の候補ロジットのみを更新し、グローバルなシーンコンテキストの進化に伴ってクロススロットのスタイル矛盾を段階的に修正する。3D-FRONTにおける実験により、最先端の家具検索性能とシーンレベルのスタイル一貫性を示し、オブジェクトレベルおよびシーンレベルの検索ベースラインよりも一貫性のある固定レイアウト家具配置を生成することを確認した。
English
Fixed-layout indoor furniture styling requires selecting assets that form a coherent room without changing the prescribed furniture categories, positions, orientations, or scales. Existing approaches typically retrieve each asset independently or rely on static local relations, making them prone to shape, material, and color conflicts after scene composition. We introduce StyleForge, a scene-level structured selection framework built on a dynamic hypergraph style field. A frozen multimodal large language model extracts structured style priors from an open-ended style request and the fixed layout, while StyleForge maintains a learnable candidate distribution for each furniture slot. Conditioned on the target style, the dynamic hypergraph style field adaptively activates and weights layout-induced hyperedges to capture higher-order dependencies among furniture. Counterfactual style preference learning then treats each candidate as a local substitution in the current style field and evaluates its contextual compatibility using Mahalanobis energies. Training alternates between optimizing the style field and the candidate logits. At inference, the model remains frozen and test-time training updates only room-specific candidate logits, progressively correcting cross-slot style conflicts as the global scene context evolves. Experiments on 3D-FRONT demonstrate state-of-the-art furniture retrieval and scene-level style coherence, producing more coherent fixed-layout furniture arrangements than object- and scene-level retrieval baselines.