ParaTempo: 시간적 신뢰도를 통한 효율적인 병렬 추론
ParaTempo: Efficient Parallel Reasoning via Temporal Confidence
August 17, 2026
저자: Xuteng Zhang, Wenhao Zeng, Xiaodong Gu, Chao Hu, Haotian Lin, Yuling Shi, Min Wang, Beijun Shen
cs.AI
초록
병렬 추론은 여러 솔루션 경로를 탐색함으로써 대규모 추론 모델의 정확성과 견고성을 향상시키지만, 그 계산 비용은 추론 깊이와 분기 수에 따라 증가한다. 이러한 병렬 경로를 관리하기 위한 기존 방법들은 일반적으로 최종 답변 합의, 국소 토큰 신뢰도, 또는 개별 중간 프로브에 의존한다. 그러나 이러한 신호들은 종종 지연되거나, 실제 추론 진행과 약하게 연관되거나, 동적 분기 수준 제어에 너무 잡음이 많다. 이러한 한계를 해결하기 위해 우리는 학습이 필요 없는 비동기 병렬 추론 프레임워크인 ParaTempo를 제안한다. ParaTempo는 답변 공간 수렴의 분기 국소 측정치인 시간적 신뢰도(temporal confidence)에 의해 구동된다. 각 분기는 잠정 답변 확률 분포를 얻기 위해 주기적으로 프로빙되며, 시간적 신뢰도는 최근 중간 프로브들이 지배적 답변에 얼마나 급격히 집중되는지를 정량화한다. 충분한 증거가 축적되면, ParaTempo는 이 단일 신호만으로 전체 제어 프로세스를 구동한다: 낮은 신뢰도의 분기는 가지치기되고, 지배적 답변에 지속적으로 고정된 분기는 조기 퇴역하며, 해제된 계산 자원은 새 분기 생성으로 재할당되고, 신뢰도 가중 투표가 집중되면 생성이 전역적으로 중단된다. 추론 궤적 간 동기화를 요구하지 않으면서, ParaTempo는 분기 수준 수렴에 기반하여 계산을 적응적으로 할당한다. 도전적인 수학 및 과학 추론 벤치마크에 대한 실험은 ParaTempo가 경쟁력 있는 정확성을 유지하면서 평균 지연 시간을 21.8-32.2%, 총 토큰 사용량을 18.1-30.3% 감소시키는 것을 보여준다. 또한, 시간적 신뢰도는 토큰 수준 및 즉각적 신호보다 미래 분기 수렴에 대해 더 강한 시간적 안정성과 예측력을 나타낸다.
English
Parallel reasoning improves the accuracy and robustness of large reasoning models by exploring multiple solution paths, but its computational cost grows with reasoning depth and branch count. Existing methods for managing these parallel paths typically rely on final-answer consensus, local token confidence, or isolated intermediate probes. However, these signals are often delayed, weakly tied to actual reasoning progress, or too noisy for dynamic, branch-level control. To address these limitations, we introduce ParaTempo, a training-free asynchronous parallel reasoning framework. ParaTempo is driven by temporal confidence, a branch-local measure of answer-space convergence. Each branch is periodically probed for a tentative answer probability distribution, and temporal confidence quantifies how sharply the recent intermediate probes concentrate on a dominant answer. Once sufficient evidence has accumulated, ParaTempo drives its entire control process from this single signal: low-confidence branches are pruned, branches that persistently commit to their dominant answer are retired early, freed computation is reallocated by forking new branches, and generation stops globally once the confidence-weighted vote concentrates. Without requiring synchronization among reasoning trajectories, ParaTempo adaptively allocates computation based on branch-level convergence. Experiments on challenging mathematical and scientific reasoning benchmarks show that ParaTempo reduces average latency by 21.8-32.2% and total token usage by 18.1-30.3% while maintaining competitive accuracy. Moreover, temporal confidence exhibits stronger temporal stability and predictive power for future branch convergence than token-level and instantaneous signals.