Flash-BoN: 확산 모델에서 추론 시간 확장을 위한 즉시 초안
Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models
July 5, 2026
저자: Ruchit Rawal, Reza Shirkavand, Sayak Paul, Yuxin Wen, Heng Huang, Yizheng Chen, Tom Goldstein, Gowthami Somepalli
cs.AI
초록
텍스트-이미지 생성을 위한 추론 시간 스케일링은 단순한 Best-of-N(BoN) 샘플링에서 중간 노이즈 제거 단계에서 후보 궤적을 검증하고 유도하는 유도 검색 방법으로 발전해 왔다. 이러한 접근법은 노이즈 제거 과정에서 언제, 얼마나 자주 검증할지에 초점을 맞추지만, 생성 자체의 비용은 대체로 고정된 것으로 간주한다. 또한, 함수 평가 횟수(NFEs)로 방법을 비교하는 표준 관행은 노이즈 제거 순방향 패스만 계산하고 검증기 오버헤드를 무시하여 효율성 순위를 왜곡할 수 있다. 실시간 평가 아래에서는 단순한 BoN이 이미 여러 유도 검색 기법과 동등하거나 더 나은 성능을 보여, 반복적인 중간 검증보다 폭넓은 탐색에 연산 자원을 투자하는 것이 더 효과적임을 시사한다. 이는 Flash-BoN으로 이어지는데, Flash-BoN은 타임스텝 절단, 레이어 스킵, 활성화 프록시라는 세 가지 상보적 가속 조정 장치를 결합하여 모델당 한 번 최적화된 단일 구성으로 저렴한 드래프트 후보의 대규모 풀을 생성한다. 이후 효율적인 다단계 검증 절차를 통해 가장 유망한 드래프트를 식별하고, 이를 전체 품질로 정제한다. 세 가지 벤치마크와 세 가지 모델 규모에서 Flash-BoN은 고정된 실시간 예산 아래 모든 기준선을 일관되게 능가하며, 그 이점은 더 큰 모델 규모에서 증가한다(+8% AUC). 또한, 본 전략이 반영 기반 프롬프트 최적화(+16% AUC)와 같은 기존의 직교 기술과 잘 결합되어 개선됨을 추가로 보여준다. 이러한 이점은 후보 다양성 증가와 상관관계가 있으며, 이는 또한 드래프트 유도 선택을 통해 강화학습 사후 학습 수렴을 가속화할 수 있게 한다.
English
Inference-time scaling for text-to-image generation has progressed from simple Best-of-N (BoN) sampling to guided search methods that verify and steer candidate trajectories at intermediate denoising steps. These approaches focus on when and how often to verify during denoising but largely treat the cost of generation itself as fixed. Moreover, the standard practice of comparing methods by number of function evaluations (NFEs) counts only denoising forward passes and ignores verifier overhead, which can distort efficiency rankings. We show that under wall-clock evaluation, simple BoN already matches or outperforms several guided search techniques, suggesting that compute is better spent on broader exploration than on repeated intermediate verification. This motivates Flash-BoN, which generates a large pool of inexpensive draft candidates by combining three complementary acceleration knobs: timestep truncation, layer skipping, and activation proxies into a single configuration optimized once per model. An efficient multi-stage verification procedure then identifies the most promising draft, which is refined at full quality. Across three benchmarks and three model scales, Flash-BoN consistently outperforms all baselines under fixed wall-clock budgets, with gains that grow at larger model scales (+8% AUC). We further show that our strategy combines well and improves existing orthogonal techniques such as reflection-based prompt optimization (+16% AUC). The gains correlate with increased candidate diversity, which also enables draft-guided selection to accelerate RL post-training convergence.