ChatPaper.aiChatPaper

보상 모델이 얼마나 빠르게 점수를 매길 수 있는가? RLHF를 위한 C++ 및 PyTorch 추론 런타임에 대한 시스템 연구

How Fast Can Reward Models Score? A Systems Study of C++ and PyTorch Inference Runtimes for RLHF

July 22, 2026
저자: Venkata Naga Sai Vishnu Rohit Pulipaka, Anish Katta, Deva Rohit Reddy Peddireddy
cs.AI

초록

RLHF 파이프라인에서 보상 점수(reward scoring)는 정책 업데이트를 차단한다. 느린 점수 계산이 전체 루프의 병목이 되는데, 모든 롤아웃이 점수를 받을 때까지 업데이트가 실행되지 않기 때문이다. 그런데도 대부분의 설정은 그냥 PyTorch eager 모드나 torch.compile을 기본값으로 사용하며, 실제로 가장 빠른 방식인지 확인하지 않는다. 점수 계산 자체는 작은 작업이다. 롤아웃 생성이 일반적인 RLHF 단계에서 훨씬 더 많은 자원을 소비한다. 그러나 점수 계산과 생성은 동일한 CPU 및 GPU 자원을 두고 경쟁하므로, 더 빠른 점수 계산 엔진이 자체적으로 단계 시간을 줄여주지는 않는다. 주로 생성이 대신 사용할 수 있는 용량을 확보해줄 뿐이다. 우리는 ONNX Runtime 기반의 네이티브 C++ 추론 엔진을 구축했다. 첫 번째 단계는 정확성 확인이었다. 출력이 CPU에서 5.7×10⁻⁶, GPU에서 4.2×10⁻³의 오차로 PyTorch 참조값과 일치하여 신뢰할 수 있는 수준이었다. 그런 다음 PyTorch eager 모드, torch.compile, FastAPI와 비교하여 CPU와 GPU 모두에서 테스트했다. CPU에서는 확실한 결과가 나왔다. 우리 엔진이 모든 기준선을 능가했으며, 신뢰 구간조차 겹치지 않았다. GPU에서는 다른 양상이 나타났다. PyTorch와 FastAPI는 이겼지만, torch.compile이 앞섰다. 추가 테스트를 통해 속도 향상의 원인이 C++ 언어 자체가 아니라 ONNX Runtime에 있음을 확인했다. 또한 배치 전략이 언어나 런타임 선택보다 더 중요했으며, 그 영향은 예상보다 컸다. 결과는 반복적이고 독립적인 실행을 통해 얻은 것이며, 단일 실행만으로는 신뢰할 만한 결과를 얻을 수 없기 때문이다.
English
In RLHF pipelines, reward scoring blocks policy updates. Slow scoring bottlenecks the entire loop, since no update runs until every rollout gets a score. And yet most setups just default to PyTorch eager mode or torch.compile, no one checks if that's actually fastest. Scoring itself is small. Rollout generation eats far more of a typical RLHF step. But scoring and generation fight over the same CPU and GPU resources, so a faster scoring engine doesn't shrink step time on its own. It mainly frees up capacity generation can use instead. We built a native C++ inference engine on ONNX Runtime. First step: confirm correctness. Output matched the PyTorch reference to 5.7 x 10^-6 on CPU and 4.2 x 10^-3 on GPU, close enough to trust. Then we tested it against PyTorch eager mode, torch.compile, and FastAPI, on both CPU and GPU. CPU was decisive. Our engine beat every baseline, confidence intervals didn't even overlap. GPU gave a different view: we beat PyTorch and FastAPI, but torch.compile came out ahead. Further testing traced the speedup to ONNX Runtime itself, not C++ as a language. And batching strategy mattered more than either the language or the runtime choice, more than we expected. The results are from repeated, independent runs, since single runs just aren't reliable enough to trust.