LongWoF-Bench:評估EvoMap基因於可驗證長工作流程任務之基準
LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks
August 24, 2026
作者: Xiao Zhang, Qumeng Sun, Jihao Li, Yiming Ren, Xiang Liu, Haoyang Zhang, Junjie Wang
cs.AI
摘要
大型語言模型日益被期望執行複雜的工作流程,其成功取決於維持相互依存的約束條件,並產出能通過嚴格端對端驗證的成果。然而,成功的執行經驗通常在一次運行後即告遺失,迫使後續模型從頭重新發掘策略與失敗模式。我們研究此類經驗是否能透過 EvoMap 被外部化並重複使用——在該框架中,經驗證器確認的執行軌跡會被整合成結構化的 Gene。為評估此設定,我們引入了長工作流程基準測試(LongWoF-Bench),包含 778 項可機器驗證的任務,涵蓋程式碼生成、代理-環境合成、數學推理與規則遵循等範疇。在 252 項具備驗證器確認之 Opus 軌跡的任務中,經過演化產生的 EvoMap Gene 在所有七個受評模型中皆優於 Skill,領先幅度為 8.7 至 15.5 個百分點,且此優勢延伸至不同模型家族中的消費級模型。相較之下,經參考蒸餾所得的 Gene 並未展現相同優勢,這表明僅有緊湊的表示形式並不足夠,Gene 的效用與經驗證的經驗來源密切相關。對於 Claude Opus,Gene 重用亦比 Skill 多完成了 39 項任務,同時將求解階段的 token 消耗降低了 9.9%。綜合而言,這些結果顯示經驗證的執行經驗可以被保留並作為可重用的外部資源加以共享,使模型能夠在不重複支付完整經驗發現成本的情況下,提升長工作流程的完成表現。
English
Large language models are increasingly expected to execute complex workflows whose success depends on maintaining interdependent constraints and producing artifacts that satisfy strict end-to-end verification. Yet successful execution experience is typically lost after a single run, forcing subsequent models to rediscover strategies and failure modes from scratch. We study whether such experience can instead be externalized and reused through EvoMap, where verifier-confirmed execution trajectories are consolidated into structured Gene. To evaluate this setting, we introduce the Long-Workflow Benchmark (LongWoF-Bench), comprising 778 machine-verifiable tasks across code generation, agent-environment synthesis, mathematical reasoning, and rule following. On the 252 tasks with verifier-confirmed Opus trajectories, evolved EvoMap Gene outperform Skill across all seven evaluated models by 8.7-15.5 percentage points, with the gains extending to consumer models from different model families. In contrast, reference-distilled Gene do not exhibit the same advantage, indicating that compact representation alone is insufficient and that Gene utility is closely associated with verified experience provenance. For Claude Opus, Gene reuse also completes 39 more tasks than Skill while reducing solve-time token consumption by 9.9%. Together, these results show that verified execution experience can be retained and shared as a reusable external resource, enabling models to improve long-workflow completion without repeatedly paying the full cost of experience discovery.