ChatPaper.aiChatPaper

LongWoF-Bench: 검증 가능한 장기 워크플로우 작업을 위한 EvoMap 유전자 평가

LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks

August 24, 2026
저자: Xiao Zhang, Qumeng Sun, Jihao Li, Yiming Ren, Xiang Liu, Haoyang Zhang, Junjie Wang
cs.AI

초록

대규모 언어 모델은 점차 상호 의존적 제약 조건을 유지하고 엄격한 종단 간 검증을 충족하는 산출물을 생성해야 하는 복잡한 워크플로우를 실행할 것으로 요구받고 있다. 그러나 성공적인 실행 경험은 일반적으로 단일 실행 후에 소실되어, 이후의 모델들은 전략과 실패 양식을 처음부터 다시 발견해야 한다. 우리는 이러한 경험이 EvoMap을 통해 외부화되고 재사용될 수 있는지 연구한다. EvoMap에서는 검증기(verifier)가 확인한 실행 궤적이 구조화된 Gene으로 통합된다. 이 설정을 평가하기 위해, 우리는 코드 생성, 에이전트-환경 합성, 수학적 추론, 규칙 준수 분야에 걸친 778개의 기계 검증 가능한 작업으로 구성된 Long-Workflow 벤치마크(LongWoF-Bench)를 도입한다. 검증기가 확인한 Opus 실행 궤적을 갖춘 252개 작업에서, 진화된 EvoMap Gene은 평가된 7개 모델 모두에 걸쳐 Skill을 8.7~15.5퍼센트 포인트 능가했으며, 이러한 이점은 서로 다른 모델 계열의 소비자용 모델에도 확장되었다. 반면, 참조 증류된 Gene은 동일한 이점을 보이지 않았는데, 이는 간결한 표현만으로는 충분하지 않으며 Gene의 유용성이 검증된 경험의 기원과 밀접하게 연관되어 있음을 시사한다. 또한 Claude Opus의 경우, Gene 재사용은 Skill보다 39개 더 많은 작업을 완료하면서 해결 시간 동안의 토큰 소비를 9.9% 줄였다. 종합하면, 이러한 결과는 검증된 실행 경험이 재사용 가능한 외부 자원으로 유지되고 공유될 수 있음을 보여주며, 모델이 경험 발견의 전체 비용을 반복적으로 지불하지 않고도 긴 워크플로우 완료를 개선할 수 있게 한다.
English
Large language models are increasingly expected to execute complex workflows whose success depends on maintaining interdependent constraints and producing artifacts that satisfy strict end-to-end verification. Yet successful execution experience is typically lost after a single run, forcing subsequent models to rediscover strategies and failure modes from scratch. We study whether such experience can instead be externalized and reused through EvoMap, where verifier-confirmed execution trajectories are consolidated into structured Gene. To evaluate this setting, we introduce the Long-Workflow Benchmark (LongWoF-Bench), comprising 778 machine-verifiable tasks across code generation, agent-environment synthesis, mathematical reasoning, and rule following. On the 252 tasks with verifier-confirmed Opus trajectories, evolved EvoMap Gene outperform Skill across all seven evaluated models by 8.7-15.5 percentage points, with the gains extending to consumer models from different model families. In contrast, reference-distilled Gene do not exhibit the same advantage, indicating that compact representation alone is insufficient and that Gene utility is closely associated with verified experience provenance. For Claude Opus, Gene reuse also completes 39 more tasks than Skill while reducing solve-time token consumption by 9.9%. Together, these results show that verified execution experience can be retained and shared as a reusable external resource, enabling models to improve long-workflow completion without repeatedly paying the full cost of experience discovery.