LongWoF-Bench: 検証可能な長期ワークフロータスクにおけるEvoMap遺伝子の評価
LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks
August 24, 2026
著者: Xiao Zhang, Qumeng Sun, Jihao Li, Yiming Ren, Xiang Liu, Haoyang Zhang, Junjie Wang
cs.AI
要旨
大規模言語モデルは、相互依存する制約を維持し、厳格なエンドツーエンド検証を満たす成果物を生成することに成功が依存する複雑なワークフローを実行することがますます期待されている。しかし、成功した実行経験は通常、一度の実行で失われ、後続のモデルは戦略や失敗モードをゼロから再発見せざるを得ない。我々は、検証器で確認された実行軌跡が構造化されたGeneに統合されるEvoMapを通じて、そのような経験を外部化して再利用できるかどうかを研究する。この設定を評価するために、コード生成、エージェント環境合成、数学的推論、およびルール遵守にわたる778の機械検証可能なタスクからなるLong-Workflow Benchmark(LongWoF-Bench)を導入する。検証器で確認されたOpus軌跡を持つ252のタスクでは、進化させたEvoMap Geneは、評価した全7モデルにおいてSkillを8.7〜15.5パーセントポイント上回り、その利得は異なるモデルファミリーの消費者向けモデルにも及ぶ。対照的に、参照蒸留されたGeneは同じ利点を示さず、コンパクトな表現だけでは不十分であり、Geneの有用性が検証された経験の由来と密接に関連していることを示している。Claude Opusでは、Gene再利用により、Skillよりも39件多いタスクを完了し、解決時のトークン消費を9.9%削減する。これらの結果は、検証済みの実行経験が再利用可能な外部リソースとして保持・共有され、モデルが経験発見の全コストを繰り返し支払うことなく、長期ワークフローの完了を改善できることを示している。
English
Large language models are increasingly expected to execute complex workflows whose success depends on maintaining interdependent constraints and producing artifacts that satisfy strict end-to-end verification. Yet successful execution experience is typically lost after a single run, forcing subsequent models to rediscover strategies and failure modes from scratch. We study whether such experience can instead be externalized and reused through EvoMap, where verifier-confirmed execution trajectories are consolidated into structured Gene. To evaluate this setting, we introduce the Long-Workflow Benchmark (LongWoF-Bench), comprising 778 machine-verifiable tasks across code generation, agent-environment synthesis, mathematical reasoning, and rule following. On the 252 tasks with verifier-confirmed Opus trajectories, evolved EvoMap Gene outperform Skill across all seven evaluated models by 8.7-15.5 percentage points, with the gains extending to consumer models from different model families. In contrast, reference-distilled Gene do not exhibit the same advantage, indicating that compact representation alone is insufficient and that Gene utility is closely associated with verified experience provenance. For Claude Opus, Gene reuse also completes 39 more tasks than Skill while reducing solve-time token consumption by 9.9%. Together, these results show that verified execution experience can be retained and shared as a reusable external resource, enabling models to improve long-workflow completion without repeatedly paying the full cost of experience discovery.