획득, 복구, 보존: 소형 모델 대화형 게임 에이전트를 위한 진단 기반 사후 훈련 레시피
Acquire, Repair, Preserve: A Diagnosis-Guided Post-Training Recipe for Small-Model Dialogue Game Agents
August 28, 2026
저자: Nan Li
cs.AI
초록
대화형 게임은 정적 벤치마크가 대체로 암묵적으로만 평가하는 능력을 시험한다. 즉, 모델은 턴에 걸쳐 상태를 유지하고, 피드백을 해석하며, 변화하는 제약 조건 하에서 유효한 행동을 선택해야 한다. 우리는 LM Playschool Challenge에서 2B 오픈 가중치 모델을 사용하여 이러한 설정을 연구했으며, 많은 실패가 단순한 광범위한 지식 부족뿐 아니라 국소적 의사결정 실패, 즉 반복된 추측, 잘못된 형식의 행동, 그리고 모델이 방금 확인한 피드백을 위반하는 행위에서 비롯됨을 발견했다. 이러한 진단은 세 단계로 구성된 훈련 방식을 정당화한다: 지도 미세 조정을 통한 광범위한 게임 참여 획득, 특정 대화 게임 계열 내에서 턴 단위 선호도 쌍을 활용한 기계적으로 검증 가능한 실패의 교정, 그리고 이러한 대화 게임을 넘어선 일반 능력의 보존. 공식 최종 평가에서 우리의 제출물은 공개 clemscore를 10.67에서 38.92로, 비공개 인도메인 점수를 13.41에서 41.17로 개선했으며, 집계 기준 정적 성능은 대략 유지했다(기준 모델 44.24 대비 44.14). 아웃오브도메인 clemscore는 7.88로 여전히 낮았으며, 가장 큰 개선은 대상 계열의 보지 못한 변형에서 집중적으로 나타났다. 우리의 결과는 광범위한 SFT가 모델 성능 향상의 대부분을 가져오며, 실패 탐지가 정밀할 때 턴 단위 국소적 감독이 효과적일 수 있고, 관찰된 전이는 주로 계열 내부에 집중됨을 시사한다.
English
Interactive dialogue games test a capability that static benchmarks largely leave implicit: a model must carry state across turns, interpret feedback, and choose valid actions under changing constraints. We study this setting in the LM Playschool Challenge with a 2B open-weight model, and find that many failures are not only broad knowledge failures but also local decision failures: repeated guesses, malformed actions, and violations of feedback that the model has just seen. These diagnostics motivate a training recipe organized around three steps: acquire broad game participation through supervised fine-tuning, repair mechanically verifiable failures within one targeted dialogue-game family using turn-local preference pairs, and preserve general capabilities beyond these dialogue games. In the official final evaluation, our submission improves public clemscore from 10.67 to 38.92 and closed in-domain score from 13.41 to 41.17, while approximately preserving aggregate static performance (44.14 vs. 44.24 for the baseline). Out-of-domain clemscore remains low at 7.88, with the largest gains concentrated in unseen variants of the targeted family. Our results suggest that broad SFT brings most of the model's capability improvement; turn-local supervision can be effective when failure detection is precise, with observed transfer concentrated primarily within-family.