GORGO: 교차 지역 네트워크 인식 LLM 서빙을 위한 온라인 튜닝
GORGO: Online Tuning for Cross-Region Network-Aware LLM Serving
June 30, 2026
저자: Alessio Ricci Toniolo, Rome Thorstenson, Abinaya Dinesh
cs.AI
초록
점점 더 많은 LLM 추론 서비스들이 클라이언트 요청을 전 세계에 분산된 엔진 복제본에 프록시하고 있다. 부하 분산 정책은 지연 시간 및 TTFT와 같은 지표를 최적화할 때 KV-캐시의 지역성, 복제본 부하, 가변적인 네트워크 지연 등 요소들을 함께 고려해야 한다. 그러나 기존 시스템은 비용 모델에서 이러한 요소들의 일부분만 평가하여 복제본 전체에 걸쳐 부하와 KV-캐시가 불균등하게 집중되는 현상을 초래한다. 본 논문에서는 조정 가능한 파라미터를 사용하여 네트워크 지연, 프리필 비용, 큐잉 지연을 전체적으로 반영하는 프록시 아키텍처 GORGO를 제시한다. LMSYS-Chat1M 및 WildChat-4.8M과 같은 오픈소스 채팅 데이터셋은 긴 컨텍스트와 높은 프리픽스 재사용률을 가진 데이터가 부족하기 때문에, 장문 컨텍스트의 프로덕션 메타데이터로부터 합성 데이터셋 ART-Chat-2.5M을 공개한다. ART-Chat-2.5M의 튜닝 윈도우에서 진화 전략을 통해 GORGO 정책의 파라미터를 유도하여 p95 TTFT를 직접 최적화한다. 홀드아웃 평가 윈도우에서는 튜닝을 통해 학습된 파라미터 값을 고정한 상태에서, 단순 세션 어피니티 및 프리픽스 캐시와 같은 기준 부하 분산 정책 대비 p95 TTFT를 6.9–15.5%, p95 종단 간(E2E) 지연 시간을 14.3–30.9% 개선한다. 코드와 ART-Chat-2.5M 데이터셋은 https://github.com/Arcadia-Research-Team/GORGO에서 확인할 수 있다.
English
Increasingly, LLM inference services proxy client requests to engine replicas distributed globally. Load-balancing policies must jointly account for factors including KV-cache locality, replica load, and variable network latency when optimizing for metrics like latency and TTFT. However, existing systems only evaluate a subset of these factors in their cost model, leading to uneven concentrations of load and KV-cache across replicas. We present GORGO, a proxy architecture that holistically factors network latency, prefill cost, and queueing delay using tunable parameters. Since open-source chat datasets such as LMSYS-Chat1M and WildChat-4.8M lack long-context, high prefix-reuse data, we release a synthetic dataset, ART-Chat-2.5M, from long-context production metadata. On a tuning window from ART-Chat-2.5M, evolutionary strategies guide the GORGO policy's parameters to directly optimize p95 TTFT. During held-out evaluation windows, we fix the parameter values learned from tuning and improve p95 TTFT by 6.9-15.5% and p95 end-to-end (E2E) latency by 14.3-30.9% over baseline load-balancing policies such as simple session affinity and prefix-cache. The code and ART-Chat-2.5M dataset can be found at https://github.com/Arcadia-Research-Team/GORGO.