ChatPaper.aiChatPaper

LongWoF-Bench:评估可验证长工作流任务中的EvoMap基因

LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks

August 24, 2026
作者: Xiao Zhang, Qumeng Sun, Jihao Li, Yiming Ren, Xiang Liu, Haoyang Zhang, Junjie Wang
cs.AI

摘要

大语言模型日益被期望执行复杂工作流,其成功取决于维护相互依赖的约束条件,并生成能够通过严格端到端验证的产物。然而,成功的执行经验通常在一次运行后即告丢失,迫使后续模型从头重新发现策略和失败模式。我们研究此类经验能否通过EvoMap实现外部化与复用,在该框架中,经验证器确认的执行轨迹被整合为结构化的Gene。为评估这一设定,我们引入了长工作流基准(LongWoF-Bench),包含778项可机器验证的任务,涵盖代码生成、智能体-环境综合、数学推理和规则遵循。在252项具有验证器确认的Opus轨迹的任务上,进化得到的EvoMap Gene在所有七个评估模型上均优于Skill,领先8.7至15.5个百分点,且收益可延伸至来自不同模型家族的消费级模型。相比之下,参考蒸馏的Gene未表现出相同优势,这表明仅靠紧凑表示并不足够,Gene的效用与经过验证的经验来源密切相关。对于Claude Opus,Gene复用还比Skill多完成了39项任务,同时将求解阶段的token消耗降低了9.9%。综合而言,这些结果表明,经过验证的执行经验可以被保留并作为可复用的外部资源进行共享,使模型能够在无需反复支付经验发现的全部代价的情况下提升长工作流的完成能力。
English
Large language models are increasingly expected to execute complex workflows whose success depends on maintaining interdependent constraints and producing artifacts that satisfy strict end-to-end verification. Yet successful execution experience is typically lost after a single run, forcing subsequent models to rediscover strategies and failure modes from scratch. We study whether such experience can instead be externalized and reused through EvoMap, where verifier-confirmed execution trajectories are consolidated into structured Gene. To evaluate this setting, we introduce the Long-Workflow Benchmark (LongWoF-Bench), comprising 778 machine-verifiable tasks across code generation, agent-environment synthesis, mathematical reasoning, and rule following. On the 252 tasks with verifier-confirmed Opus trajectories, evolved EvoMap Gene outperform Skill across all seven evaluated models by 8.7-15.5 percentage points, with the gains extending to consumer models from different model families. In contrast, reference-distilled Gene do not exhibit the same advantage, indicating that compact representation alone is insufficient and that Gene utility is closely associated with verified experience provenance. For Claude Opus, Gene reuse also completes 39 more tasks than Skill while reducing solve-time token consumption by 9.9%. Together, these results show that verified execution experience can be retained and shared as a reusable external resource, enabling models to improve long-workflow completion without repeatedly paying the full cost of experience discovery.