ChatPaper.aiChatPaper

AgentMercury: 大規模なビジネスシナリオ向けに検証可能な環境を合成できるエージェント

AgentMercury: Your Agent Can Synthesize Verifiable Environments for Business Scenarios at scale

August 21, 2026
著者: Minbyul Jeong, Chanwoong Yoon
cs.AI

要旨

エージェントは環境との相互作用を通じて行動を学習するが、訓練に使用される環境は多くの場合、事前定義されたタスクやベンチマークを中心に手作業で構築または合成される。このタスク中心のパラダイムでは、多様なタスクが基盤となる世界から自然に創発し得る現実的かつ進化的なワークフローを反映した環境をスケーラブルに構築することが困難である。本稿では、高レベルのビジネスシナリオから実行可能な環境を合成するためのスケーラブルなフレームワークであるAgentMercuryを提案する。AgentMercuryは特定のタスクのために環境を構築するのではなく、まずエンティティ、サービス、ツール、状態、および実行可能なサービス間不変条件を備えた永続的な世界をインスタンス化し、そこから多様なタスクと相互作用の軌跡が続いて創発することを可能にする。我々は14産業・50カ国にわたる4,783の実行可能な環境を構築し、それらを強化学習の訓練基盤として使用した。評価ベンチマークを対象とせずに生成されたにもかかわらず、これらのビジネス指向環境で訓練された方策は、エンタープライズワークフローと、推論・コーディング・科学計算・ツール使用を網羅する領域外ベンチマークの両方において大幅に向上した。実験では、Qwen3.5-4BはAgentMercury環境での訓練後に、EnterpriseOps-GYMにおいて12.3から15.7へ、AIME26において45.9から56.0へ改善した。さらに、構築プロセス自体が学習可能であることを示す。すなわち、Qwen3.5-35B-A3Bを構築トレースでファインチューニングすることにより、未見のビジネスシナリオにおける実行可能な世界の作成成功率が3.3%から83.3%に向上した。これらの結果は、シナリオに基盤を置く環境が、ベンチマーク特化型の訓練を超えた有用かつ一般化可能な学習シグナルを提供できること、またその構築自体が学習可能な能力となり得ることを示している。
English
Agents learn to act through interaction with environments, yet the environments used for training are often manually constructed or synthesized around predefined tasks and benchmarks. This task-centric paradigm makes it difficult to scale environments that reflect realistic and evolving workflows where diverse tasks can naturally emerge from the underlying world. We introduce AgentMercury, a scalable framework for synthesizing executable environments from high-level business scenarios. Rather than constructing an environment for a specific task, AgentMercury first instantiates a persistent world with entities, services, tools, state, and executable cross-service invariants, from which diverse tasks and interaction trajectories can subsequently emerge. We construct 4,783 executable environments spanning 14 industries and 50 countries, and use them as training substrates for reinforcement learning. Despite being generated without targeting the evaluation benchmarks, policies trained on these business-oriented environments improve substantially on both enterprise workflows and out-of-domain benchmarks spanning reasoning, coding, scientific computing, and tool use. In our experiments, Qwen3.5-4B improves from 12.3 to 15.7 on EnterpriseOps-GYM and from 45.9 to 56.0 on AIME26 after training on AgentMercury environments. We further show that the construction process itself can be learned: fine-tuning Qwen3.5-35B-A3B on construction traces increases executable-world authoring success from 3.3% to 83.3% on held-out business scenarios. These results show that scenario-grounded environments can provide useful and generalizable learning signals beyond benchmark-specific training, while their construction can itself become a learnable capability.