양자화 인지 치유: 압축된 4비트 LLM 복구를 위한 실용적 방법
Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs
August 21, 2026
저자: Bakbergen Ryskulov, Iker García-Ferrero, David Montero, David Jansen, Ali Hashemi, Jezabel R. Garcia, Antonio Tiene, Román Orús
cs.AI
초록
대규모 언어 모델을 저비용으로 서빙하는 것은 점점 더 파라미터의 일부로 구조적으로 압축되고 4비트로 양자화된 모델을 배포하는 것을 의미한다. 이러한 단계는 함께 적용될 때 추론, 수학, 코딩, 긴 문맥(long-context) 동작을 충분히 저하시켜 배포 전에 회복(힐링) 단계를 요구한다. 기본 레시피인 양자화 인식 훈련(QAT)은 압축되고 양자화된 모델을 하드 레이블에 재적합시키지만, 우리의 파이프라인에서는 QAT가 천천히 수렴하다가 최고점을 지나 붕괴했다. 우리는 대신 양자화 인식 힐링(QAH)을 채택했다. 구조적으로 압축된 모델은 전체 정밀도로 독립적으로 훈련된 적이 없기 때문에, 해당 bfloat16 체크포인트는 원본 모델의 증류 회복 근사치이다. QAH는 4비트 학생 모델을 원본의 비압축 모델에서 직접 증류한다. GPT-OSS 120B에서 60B로, 그리고 MXFP4로 이어지는 파이프라인에서 QAH 학생 모델은 약 4배 적은 가중치 메모리와 교사 모델 대비 절반의 파라미터 수로 9개 벤치마크 중 7개에서 bfloat16 소스 모델과 동일하거나 더 나은 성능을 보였으며, Hypernova-60B라는 이름의 오픈 가중치 모델로 공개되었다. 대응되는 QAT 기준선과 비교하면, QAH는 약 7배 더 빠르게 유사한 최고 성능에 도달하고, 수동으로 조정된 조기 종료 없이 지속적인 훈련에서도 안정적으로 유지된다. 또한 우리는 분산 훈련 백엔드 간의 크고 재현 가능한 품질 격차를 포함한 배포 교훈을 보고한다. 우리의 목표는 수 주에 걸친 하이퍼파라미터 탐색 없이 배포할 수 있는 레시피를 제시하는 것이다.
English
Serving large language models cheaply increasingly means shipping models that are both structurally compressed to a fraction of their parameters and quantized to 4 bits. Together these steps degrade reasoning, mathematics, coding, and long-context behavior enough to require a recovery, or healing, stage before deployment. The default recipe, quantization-aware training (QAT), re-fits the compressed, quantized model to hard labels; in our pipeline it converged slowly and collapsed past its peak. We adopted Quantization-Aware Healing (QAH) instead. Because a structurally compressed model is never independently trained at full precision, its bfloat16 checkpoint is a distillation-recovered approximation of the original; QAH distills the 4-bit student directly from the original, uncompressed model. On a GPT-OSS 120B to 60B to MXFP4 pipeline, the QAH student matches or beats its bfloat16 source on 7 of 9 benchmarks at roughly 4 times less weight memory and half the teacher's parameter count, and is released open-weight as Hypernova-60B. Against a matched QAT baseline it reaches a comparable peak about 7 times faster and stays stable under continued training, without hand-tuned early stopping. We also report deployment lessons, including a large, reproducible quality gap between distributed-training backends. Our aim is a recipe deployable without a multi-week hyper-parameter search.