ChatPaper.aiChatPaper

다단계 시각적 추론을 위한 계층적 잡음 제거

Hierarchical Denoising For Multi-Step Visual Reasoning

July 16, 2026
저자: Zezhong Qian, Xiaowei Chi, Chak-Wing Mak, Tianze Zhou, Ruibin Yuan, Yuhan Rui, Hengzhe Sun, Zhuoqun Wu, Yuming Li, Siyuan Qian, Sirui Han, Shanghang Zhang
cs.AI

초록

비디오 모델은 비전 기초 모델로 진화하고 있지만, 여전히 인간과 같은 다단계 추론이 부족하다. 스트리밍 자기회귀 확산 모델은 효율적이지만 추론에 한계가 있는 반면, 양방향 확산은 밀집된 프레임 수준 잡음 제거로 인해 높은 추론 비용으로 전역 수정을 가능하게 한다. 두 패러다임 모두 복잡한 추론 작업에 대해 논리적 일관성과 저지연 스트리밍을 달성하는 데 어려움을 겪는다. 본 논문은 다단계 추론을 위해 계층적 잠재 변수를 인과적 비디오 생성에 통합하는 통합 프레임워크인 HDR(시각적 추론을 위한 계층적 잡음 제거)을 제안한다. HDR은 비디오 잠재 변수를 트리 구조 계층으로 구성하여 스트리밍 출력 전에 대략적에서 세부적인 추론을 가능하게 한다. 대략적 잡음 제거 계층은 전역 계획을 위해 불확실한 가설을 보존하고, 세부 계층은 이를 점진적으로 구체적인 시각 상태로 정제한다. 희소 계층적 주의 패턴(SHAP)은 시간적 주의 비용을 더욱 감소시킨다. 본 논문은 분포 외 사례를 포함하는 수준별 다단계 비디오 추론 벤치마크를 소개하며, 미로 탐색, 하노이 탑, 한 줄 그리기, 슬라이딩 퍼즐, 소코반, 물 따르기의 여섯 가지 작업을 포함한다. 스트리밍 자기회귀 확산 기준선과 비교하여 HDR은 성공률을 34.22에서 60.29(76.2% 상대적 향상)로 개선하고 평균 진행률을 76.00에서 89.56으로 증가시켜 더 일관된 추론 궤적을 보여준다. HDR은 잠재 변수당 0.70초의 저지연 스트리밍을 유지하여 양방향 확산보다 54.2배 빠른 추론을 달성한다. 또한 2%의 학습 데이터만으로 전체 데이터 성능의 82.9%를 유지하며, 양방향 확산의 52.0%와 대조된다. 실제 로봇 실험은 물리적 상호작용 및 세계 모델링에 대한 HDR의 잠재력을 추가로 입증한다. 프로젝트 데모: https://hierarchical-diffusion-reasoning.github.io/.
English
Video models are evolving into vision foundation models, yet they still lack human-like multi-step reasoning. Streaming autoregressive diffusion models are efficient but limited in reasoning, while bidirectional diffusion enables global revision with high inference costs due to dense frame-level denoising. Both paradigms struggle to achieve logical consistency and low-latency streaming for complex reasoning tasks. We propose HDR (Hierarchical Denoising for Visual Reasoning), a unified framework that integrates hierarchical latents into causal video generation for multi-step reasoning. HDR organizes video latents into a tree-structured hierarchy, enabling coarse-to-fine reasoning before streaming output. Coarse denoising layers preserve uncertain hypotheses for global planning, while finer layers progressively refine them into concrete visual states. A sparse hierarchical attention pattern (SHAP) further reduces temporal attention costs. We introduce a level-stratified multi-step video reasoning benchmark with out-of-distribution cases, covering six tasks: maze navigation, Tower of Hanoi, one-line drawing, sliding puzzle, Sokoban, and water pouring. Compared with streaming autoregressive diffusion baselines, HDR improves success from 34.22 to 60.29 (76.2% relative gain) and increases average progress from 76.00 to 89.56, demonstrating more consistent reasoning trajectories. HDR maintains low-latency streaming at 0.70 seconds per latent, achieving 54.2 times faster inference than bidirectional diffusion. It also retains 82.9% of full-data performance with only 2% training data, compared with 52.0% for bidirectional diffusion. Real-world robot experiments further demonstrate HDR's potential for physical interaction and world modeling. Project demo: https://hierarchical-diffusion-reasoning.github.io/.