StyleForge: 하이퍼그래프 필드에서의 반사실적 추론을 통한 실내 가구 스타일링
StyleForge: Indoor Furniture Styling by Counterfactual Reasoning in a Hypergraph Field
August 3, 2026
저자: Lingwei Dang, Shishuo Shang, Pan Liu, Jiajia Cheng, Ziyan Qiu, Zhenhao Zhang, Yufei Zhu, Shenghui Huang, Qingxin Xiao, Yun Hao, Juntong Li, Qingyao Wu
cs.AI
초록
고정 레이아웃 실내 가구 스타일링은 지정된 가구 범주, 위치, 방향, 또는 크기를 변경하지 않으면서 조화로운 방을 구성하는 에셋을 선택해야 한다. 기존 접근법은 일반적으로 각 에셋을 독립적으로 검색하거나 정적 국소 관계에 의존하기 때문에, 장면 구성 후 형태, 재질, 색상 충돌이 발생하기 쉽다. 본 논문에서는 동적 하이퍼그래프 스타일 필드(dynamic hypergraph style field) 기반의 장면 수준 구조적 선택 프레임워크인 StyleForge를 제안한다. 동결된(frozen) 멀티모달 대규모 언어 모델은 개방형 스타일 요청과 고정 레이아웃에서 구조화된 스타일 사전 정보를 추출하며, StyleForge는 각 가구 슬롯에 대해 학습 가능한 후보 분포를 유지한다. 대상 스타일을 조건으로, 동적 하이퍼그래프 스타일 필드는 레이아웃 유도 하이퍼에지를 적응적으로 활성화하고 가중치를 부여하여 가구 간 고차 의존성을 포착한다. 반사실적 스타일 선호 학습(counterfactual style preference learning)은 각 후보를 현재 스타일 필드에서의 국소적 대체로 간주하고 마할라노비스 에너지를 사용하여 맥락적 호환성을 평가한다. 훈련은 스타일 필드 최적화와 후보 로짓 최적화를 번갈아 수행한다. 추론 시, 모델은 동결된 상태를 유지하며, 테스트 시점 학습(test-time training)은 방 특정 후보 로짓만 업데이트하여 전역 장면 맥락이 진화함에 따라 슬롯 간 스타일 충돌을 점진적으로 교정한다. 3D-FRONT 실험에서 제안 방법은 최첨단 가구 검색 및 장면 수준 스타일 일관성을 보여주며, 객체 및 장면 수준 검색 베이스라인보다 더 일관된 고정 레이아웃 가구 배치를 생성한다.
English
Fixed-layout indoor furniture styling requires selecting assets that form a coherent room without changing the prescribed furniture categories, positions, orientations, or scales. Existing approaches typically retrieve each asset independently or rely on static local relations, making them prone to shape, material, and color conflicts after scene composition. We introduce StyleForge, a scene-level structured selection framework built on a dynamic hypergraph style field. A frozen multimodal large language model extracts structured style priors from an open-ended style request and the fixed layout, while StyleForge maintains a learnable candidate distribution for each furniture slot. Conditioned on the target style, the dynamic hypergraph style field adaptively activates and weights layout-induced hyperedges to capture higher-order dependencies among furniture. Counterfactual style preference learning then treats each candidate as a local substitution in the current style field and evaluates its contextual compatibility using Mahalanobis energies. Training alternates between optimizing the style field and the candidate logits. At inference, the model remains frozen and test-time training updates only room-specific candidate logits, progressively correcting cross-slot style conflicts as the global scene context evolves. Experiments on 3D-FRONT demonstrate state-of-the-art furniture retrieval and scene-level style coherence, producing more coherent fixed-layout furniture arrangements than object- and scene-level retrieval baselines.