LongStraw: 고정 GPU 예산 하에서 200만 토큰을 넘는 장문맥 강화학습

LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget

July 16, 2026
저자: Changhai Zhou, Kieran Liu, Yuhua Zhou, Qian Qiao, Jun Gao, Harry Zhang, Irvine Lu, Nolan Ho, Lucian Li, Andrew Lei, Cleon Cheng, Steven Chiang, Yihang Zeng, Di Zhang, Rio Yang, Kaijie Chen, Andrew Chen, Pony Ma, Weizhong Zhang, Cheng Jin
cs.AI

초록

추론 컨텍스트 길이와 강화학습(RL) 후속 학습(Post-training) 사이의 격차가 점점 벌어지고 있다. 즉, 추론 시스템은 백만 토큰 컨텍스트에 근접하는 반면, 후속 학습 작업은 256K 토큰 이하에 머물러 배포 시 길이 일반화(Length Generalization)에 의존하는 경우가 많다. 이러한 격차는 관찰 결과, 도구 출력, 문서 및 이전 결정이 긴 궤적에 걸쳐 축적되는 AI 에이전트에게 특히 중요하다. LongStraw는 고정 GPU 예산 하에서 백만 토큰 규모의 RL 후속 학습을 위한 아키텍처 인식 실행 스택(Architecture-aware Execution Stack)으로, 그룹 상대 정책 최적화(GRPO)를 통해 구현된다. 이는 autograd 없이 공유 프롬프트를 평가하고, 이후 토큰에 필요한 모델별 상태만 유지하며, 단시간 응답 분기(Short Response Branch)를 한 번에 하나씩 재생(Replay)하여, 추가 재생 시간을 대가로 활성 학습 그래프의 크기를 줄인다. 우리는 이를 하이브리드 순환 및 전체 주의(Hybrid Recurrent and Full-attention) 기반의 Qwen3.6-27B와 압축 주의 혼합 전문가(Compressed-attention Mixture-of-experts) 기반의 GLM-5.2에 구현했다. 8개의 H20 GPU에서 LongStraw는 그룹 크기 2와 8에 대해 210만 위치에서 그룹화된 Qwen의 스코어링 및 응답 역전파(Backward)를 완료한다. 그룹 크기를 늘려도 최대 할당 메모리는 0.21GB만 증가하며, 별도의 스트레스 테스트에서는 446만 위치에 도달한다. 32개의 H20 GPU에서는 GLM-5.2의 전체 78개 레이어에 걸쳐 210만 토큰 프롬프트에 대한 종단 간 LongStraw 실행 경로를 검증한다. 이러한 실험은 캡처된 프롬프트 상태가 분리(Detach)되고 일부 분산 포워드 및 그래디언트 구성 경로가 미완성 상태로 남아 있기 때문에, 완전한 훈련 정확성보다는 실행 용량을 입증하는 데 초점을 맞춘다.
English
A growing gap separates inference context lengths from RL post-training: inference systems are approaching million-token contexts, while post-training workloads often remain at 256K tokens or below and rely on length generalization at deployment. The gap is especially important for AI agents, whose observations, tool outputs, documents, and prior decisions accumulate over long trajectories. LongStraw is an architecture-aware execution stack for million-token RL post-training under a fixed GPU budget, instantiated with Group Relative Policy Optimization (GRPO). It evaluates the shared prompt without autograd, retains only model-specific state needed by later tokens, and replays short response branches one at a time, reducing the live training graph at the cost of additional replay time. We implement it for the hybrid recurrent and full-attention Qwen3.6-27B and the compressed-attention mixture-of-experts GLM-5.2. On eight H20 GPUs, LongStraw completes grouped Qwen scoring and response backward at 2.1M positions for groups of 2 and 8; increasing the group size adds only 0.21 GB of peak allocated memory, while a separate stress test reaches 4.46M positions. On 32 H20 GPUs, we validate the end-to-end LongStraw execution path for a 2.1M-token prompt across all 78 layers of GLM-5.2. These experiments establish execution capacity rather than complete training correctness because the captured prompt state is detached and some distributed forward and gradient composition paths remain incomplete.
PDF402July 18, 2026