ChatPaper.aiChatPaper

AgentMercury:你的智能體可大規模合成商業場景之可驗證環境

AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale

August 21, 2026
作者: Minbyul Jeong, Chanwoong Yoon
cs.AI

摘要

智能體透過與環境互動來學習行為,然而用於訓練的環境往往圍繞預定義任務和基準,以人工建構或合成方式產生。這種以任務為中心的研究範式,使得擴展能反映真實且不斷演進的工作流程之環境變得困難,因為在這樣的世界中,多樣化的任務自然會從底層湧現。我們提出AgentMercury,一個可擴展的框架,能從高階業務場景合成可執行的環境。AgentMercury並非為特定任務建構環境,而是先實例化一個持續存在的世界,其中包含實體、服務、工具、狀態以及可執行的跨服務不變量,使多樣化的任務和互動軌跡得以從中湧現。我們建構了涵蓋14個產業、50個國家的4,783個可執行環境,並將其作為強化學習的訓練基質。儘管這些環境在生成時並未針對評測基準,但在此類業務導向環境上訓練的策略,在企業工作流程以及涵蓋推理、程式設計、科學計算和工具使用的跨領域基準上均有顯著提升。在我們的實驗中,Qwen3.5-4B在AgentMercury環境上訓練後,於EnterpriseOps-GYM的成績從12.3提升至15.7,於AIME26的成績從45.9提升至56.0。我們進一步證明,建構過程本身是可以學習的:在建構軌跡上微調Qwen3.5-35B-A3B,使留出業務場景的可執行世界創作成功率從3.3%提升至83.3%。這些結果表明,以場景為基礎的環境能提供超越特定基準訓練的有用且可泛化的學習訊號,而其建構過程本身也能成為一種可學習的能力。
English
Agents learn to act through interaction with environments, yet the environments used for training are often manually constructed or synthesized around predefined tasks and benchmarks. This task-centric paradigm makes it difficult to scale environments that reflect realistic and evolving workflows where diverse tasks can naturally emerge from the underlying world. We introduce AgentMercury, a scalable framework for synthesizing executable environments from high-level business scenarios. Rather than constructing an environment for a specific task, AgentMercury first instantiates a persistent world with entities, services, tools, state, and executable cross-service invariants, from which diverse tasks and interaction trajectories can subsequently emerge. We construct 4,783 executable environments spanning 14 industries and 50 countries, and use them as training substrates for reinforcement learning. Despite being generated without targeting the evaluation benchmarks, policies trained on these business-oriented environments improve substantially on both enterprise workflows and out-of-domain benchmarks spanning reasoning, coding, scientific computing, and tool use. In our experiments, Qwen3.5-4B improves from 12.3 to 15.7 on EnterpriseOps-GYM and from 45.9 to 56.0 on AIME26 after training on AgentMercury environments. We further show that the construction process itself can be learned: fine-tuning Qwen3.5-35B-A3B on construction traces increases executable-world authoring success from 3.3% to 83.3% on held-out business scenarios. These results show that scenario-grounded environments can provide useful and generalizable learning signals beyond benchmark-specific training, while their construction can itself become a learnable capability.