ChatPaper.aiChatPaper

FACET:在終端任務合成中保留來源意圖與可執行狀態

FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis

August 19, 2026
作者: Kou Shi, Zun Wang, Qisheng Su, Shiting Huang, Ziao Zhang, Zhen Fang, Qingnan Ren, Jin Liu, Yu Zeng, Yiming Zhao, Lin Chen, Zehui Chen, Feng Zhao
cs.AI

摘要

訓練終端智能體需要可擴展的可執行監督,然而合成高品質的終端任務仍然具有挑戰性。每個任務都耦合了指令、初始化環境、參考解答和可執行的驗證器;如果這些工件是基於不一致的假設生成的,最終任務可能無法解決或遭到錯誤評估。與此同時,多階段合成可能會丟棄原始來源中所編碼的目標、依賴關係、狀態轉換和程序性約束。我們提出 FACET(Fine-grained Agentic Construction of Executable Tasks,可執行任務的細粒度智能體式建構),這是一個同時處理資訊保存與跨工件一致性的框架。FACET 將相關的智能體技能重構為連貫且資訊豐富的情境,然後在生成最終任務工件之前實現並修復執行環境。最終的容器狀態可作為指令、解答和驗證器的共享基礎,而以執行為基礎的驗證與針對性修復則能修正特定工件的失敗,無須不必要地重新生成有效元件。FACET 產生了具有密集可執行檢查的複雜終端任務,而從這些任務中收集的成功軌跡提供了有效且資料效率高的監督。跨多種規模微調模型持續提升了在 Terminal-Bench 2.1 上的效能,而對替代生成方案的分析則支持了以環境為基礎的建構對於任務有效性及解答與驗證器一致性的重要性。這些結果確立了原始意圖保留與共享可執行狀態基礎,作為可擴展終端任務合成的關鍵原則。
English
Training terminal agents requires scalable executable supervision, yet synthesizing high-quality terminal tasks remains challenging. Each task couples an instruction, an initialized environment, a reference solution, and an executable verifier; if these artifacts are generated from inconsistent assumptions, the resulting task may be unsolvable or incorrectly evaluated. Meanwhile, multi-stage synthesis can discard the goals, dependencies, state transitions, and procedural constraints encoded in the original sources. We present FACET (Fine-grained Agentic Construction of Executable Tasks), a framework that addresses both information preservation and cross-artifact consistency. FACET reconstructs related agent skills into coherent, information-rich scenarios, then realizes and repairs the execution environment before generating the final task artifacts. The resulting container state serves as shared grounding for the instruction, solution, and verifier, while execution-based validation and targeted repair correct artifact-specific failures without unnecessarily regenerating valid components. FACET produces complex terminal tasks with dense executable checks, and successful trajectories collected from these tasks provide effective, data-efficient supervision. Fine-tuning models across multiple scales consistently improves performance on Terminal-Bench 2.1, while analyses of alternative generation schemes support the importance of environment-grounded construction for task validity and solution-verifier alignment. These results establish source-intent preservation and shared executable-state grounding as key principles for scalable terminal-task synthesis.