xHC: 확장된 하이퍼 연결
xHC: Expanded Hyper-Connections
July 16, 2026
저자: Xiangdong Zhang, Xiaohan Qin, Sunan Zou, Tuo Dai, Xiaoming Shi, Huaijin Wu, Yebin Yang, Zhuo Xia, Shaofeng Zhang, Lin Yao, Yuliang Liu, Yu Cheng, Junchi Yan
cs.AI
초록
하이퍼 연결(HC)은 트랜스포머의 잔차 스트림을 N개의 병렬 스트림으로 확장하여 모델 폭과 깊이를 넘어서는 메모리 스케일링을 제공합니다. 다양체 제약 HC(mHC)는 이러한 공식을 대규모에서 안정화합니다. N=1에서 N=4까지의 큰 성능 향상은 잔차 스트림 확장이 유망한 스케일링 축임을 시사합니다. 그러나 기존 HC 계열 방법은 일반적으로 N=4에서 멈춥니다. 실험 결과 그 이유가 밝혀졌습니다. 이 지점을 넘어 mHC를 확장하면 성능 향상이 감소하고 훈련 비용이 급격히 증가합니다. 이러한 한계는 두 가지 병목 현상, 즉 증가하는 스트림 수에 대한 불충분한 쓰기-백 정보와 N에 따라 비용이 세제곱으로 증가하는 잔차 혼합 생성에 기인합니다. 두 병목 현상을 해결하기 위해, 우리는 N=4를 넘어 유의미한 확장을 달성한 최초의 HC 계열 방법인 xHC(확장된 하이퍼 연결)를 제안합니다. xHC는 더 풍부한 쓰기-백을 위한 시간적 특징 증강과 전체 잔차 상태에 대한 밀집 접근을 유지하면서 N=16개의 스트림 중 k=4개만 업데이트하는 희소 잔차 스트림 구조를 결합합니다. 18B 및 28B MoE 모델 전반에 걸쳐 xHC는 강력하고 일관된 하위 작업 성능 개선을 제공합니다. 18B MoE 모델에서 xHC는 mHC 대비 평균 하위 작업 점수를 4.0포인트 향상시키는 동시에 기본(vanilla) 기준선 대비 훈련 FLOPs를 약간만 증가시킵니다. 스케일링 법칙 실험에 따르면 vanilla과 mHC는 동일한 손실에 도달하기 위해 각각 xHC 대비 1.50배 및 1.19배의 연산이 필요합니다. 실용적인 대규모 N 훈련을 위해서는 확장된 잔차 상태로 인한 메모리 트래픽을 제어해야 합니다. 따라서 우리는 xHC-Flash를 도입합니다. 이는 전체 xHC의 이점을 유지하면서 서브레이어당 메모리 트래픽을 73.5C에서 40C로 줄이며, 이는 N=4에서 mHC가 필요로 하는 34C와 유사한 수준입니다. xHC와 xHC-Flash는 함께 대규모 N 잔차 스트림 확장을 LLM 사전학습에 효과적이고 실용적으로 만듭니다.
English
Hyper-Connections (HC) expand the residual stream of Transformers into N parallel streams, providing a form of memory scaling beyond model width and depth. Manifold-Constrained HC (mHC) stabilizes this formulation at scale. The large gains from N{=}1 to N{=}4 suggest residual-stream expansion as a promising scaling axis. However, existing HC-family methods typically stop at N{=}4. Our experiments reveal why: scaling mHC beyond this point yields diminishing performance gains and rapidly increasing training cost. We attribute this limitation to two bottlenecks: insufficient write-back information for an expanding number of streams and residual-mixing generation whose cost scales cubically with N. To address both bottlenecks, we propose xHC (Expanded Hyper-Connections), the first HC-family method to achieve meaningful expansion beyond N{=}4. xHC combines temporal feature augmentation for richer write-back with a sparse residual-stream architecture that updates only k=4 of the N=16 streams while retaining dense access to the full residual state. Across 18B and 28B MoE models, xHC delivers strong and consistent downstream improvements. On an 18B MoE model, xHC improves the average downstream score by 4.0 points over mHC, while adding only modest training FLOPs over the vanilla baseline. Scaling-law experiments show that the vanilla and mHC require 1.50times and 1.19times the compute of xHC, respectively, to reach the same loss. Practical large-N training also requires controlling memory traffic from the expanded residual state. We therefore introduce xHC-Flash, which reduces the per-sublayer memory traffic from 73.5C to 40C, comparable to the 34C required by mHC at N{=}4, while retaining the gains of full xHC. Together, xHC and xHC-Flash make large-N residual-stream expansion effective and practical for LLM pre-training.