ChatPaper.aiChatPaper

한 번의 성공은 신뢰성이 아니다: 상태 기반 비즈니스 워크플로우 에이전트를 위한 샌드박스이자 벤치마크인 Thinkingbox

One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows

August 20, 2026
저자: Zhuochun Li, Youngmin Ko, Ali Keramati, Nicola Ferri, Susana Palmaz Lopez Pelaez, Liang-Chun Tsai, Calvin Wang, Mirco Milletari, Tuhin Kundu, Vadim Smolyakov, Kjartan Olafsson, Tommy Guy
cs.AI

초록

최근 에이전트 벤치마크는 코드 수리에서 웹 탐색, 앱 API, 함수 호출에 이르기까지 평가를 점점 더 실행 가능한 환경에 기반을 두고 있다. 그러나 코드를 넘어선 중대한 작업을 완료하는 것은 그럴듯한 응답이나 유효한 도구 호출을 생성하는 것 이상을 요구한다. 에이전트는 여러 턴에 걸쳐 누락된 정보를 수집하고, 도메인 정책을 따르고, 의존 관계에 있는 도구들을 조정하며, 부수 효과 없이 올바른 지속 상태 전이를 실현해야 한다. 본 논문에서는 도구-에이전트-사용자 상호작용을 위한 샌드박스인 Thinkingbox를 소개한다. 이 샌드박스는 격리된 MCP 호환 도구 세션, 완전한 실행 추적, 그리고 최종 백엔드 상태에 대한 결과 평가를 제공한다. 이 샌드박스를 기반으로 구축된 Thinkingbox-bench는 소매, 호텔·접객, 자동차 보험, 네오뱅크 내부 IT, 컨설팅 IT/HR 지원 등 다양한 시나리오에 걸쳐 507개의 정책 조건부 워크플로를 포함한다. 각 시도는 유효한 궤적은 수용하면서 잘못되거나 누락되었거나 추가된 효과는 거부하는 작업별 실행 가능 검사로 평가되며, 지정된 작업은 최종 응답의 필수 속성도 추가로 검사한다. 독점 모델과 공개 가중치 모델을 통틀어 가장 강력한 모델은 pass@1에서 65.36%를 달성하지만, pass^20에서는 25.25%에 그친다. 더욱이 많은 실패한 시도는 깨끗한 종료와 유효한 상태 변경 작업을 보여주는데, 이는 응답이나 도구 호출 수준의 신호가 종단 간 작업 완료의 명확한 대리 지표가 아님을 시사한다. Thinkingbox-bench는 성공적인 궤적을 우연히 발견하는 것과 상태 저장 비즈니스 작업을 안정적으로 완료하는 것 사이에 큰 격차가 있음을 드러낸다. 우리는 Thinkingbox와 Thinkingbox-Bench를 모두 공개한다: https://github.com/microsoft/thinkingbox
English
Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet completing consequential work beyond code requires more than producing a plausible response or valid tool call: agents must gather missing information over multiple turns, follow domain policies, coordinate dependent tools, and realize the correct persistent state transition without collateral effects. In this paper, we introduce Thinkingbox, a sandbox for tool-agent-user interaction that provides isolated MCP-compatible tool sessions, complete execution traces, and outcome evaluation over terminal backend state. Built on this sandbox, Thinkingbox-bench contains 507 policy-conditioned workflows across numerous scenarios, including retail, hospitality, auto insurance, neobank internal IT, and consulting IT/HR support. Each attempt is evaluated by task-specific executable checks that accept valid trajectories while rejecting wrong, missing, or extra effects; designated tasks additionally check required properties of the final response. Across proprietary and open-weight models, the strongest achieves 65.36% pass@1, but only 25.25% pass^20. Moreover, many failed trials show clean termination and valid state-changing actions, showing that response or tool-call-level signals are not clear proxies for end-to-end task completion. Thinkingbox-bench reveals a large gap between occasionally finding a successful trajectory and reliably completing stateful business tasks. We release both Thinkingbox and Thinkingbox-Bench: https://github.com/microsoft/thinkingbox