Tencent WorkBuddy Bench: 오염 저항적 태스크 구성을 갖춘 다중 도메인 코딩 에이전트 벤치마크
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction
July 23, 2026
저자: Tencent WorkBuddy Bench Team, Siqi Cai, Shaopeng Chen, Xiang Fei, Yong Mao, Zihan Xu, Zhiheng Lyu, Zhijian Shao, Yuchen Shi, Shuwen Zhang, Chaofan Qiu, Linjie Che, Xiaoxi Zhao, Feng Wu, Kai Zhang, Chaofan Zhu, Yubin Qi, Xiaoyun Liang, Peijie Dong, Yunhao Zhang, Yuanjie Zhu, Ling Jiang, Xianjun Zhang, Zhehang Chu, Anyuan Sang, Zhen Feng, Sen Nie, Shi Wu, Yuanzhen Xu, Xin Li, Ning Yang, Zhiqiang Dong, Hande Dong, Qiang Lin, Yi Liu, Yunsheng Wu, Ke Li, Xing Sun
cs.AI
초록
Tencent WorkBuddy Bench를 소개합니다. 이는 코딩 에이전트를 위한 다중 도메인 평가 제품군으로, 본 보고서는 그 구축 방법론, 채점 프로토콜, 그리고 교차 모델 리더보드를 문서화합니다. 핵심에는 코드, 웹, 오피스, 보안의 네 가지 작업 도메인에 걸쳐 분포 기반 코딩 에이전트 태스크를 구성하고 실행하기 위한 통합 평가 프레임워크가 있습니다. 공개된 이슈 텍스트를 변형하는 대신, 모든 태스크는 실제 커밋, 풀 리퀘스트 또는 비즈니스 시나리오에서 역설계되어 짧고 구어체적인 역할극 요청으로 재작성되므로, 태스크 프롬프트는 기저의 이슈, 풀 리퀘스트 또는 커밋 스레드를 웹 검색으로 복구할 수 없습니다. 데이터셋이 태스크 디렉터리, 환경 이미지, 평가 도구, 테스트, 참조 솔루션 등으로 공개적으로 제공되기 때문에, 오염 저항성은 비밀이 아닌 이러한 구성 방식과 데이터셋 버전 관리에 기반합니다. 네 가지 하위 집합(리포지토리 수준 엔지니어링, 프론트엔드 개발, 오피스 및 비즈니스 워크플로, 레드-/블루팀 보안)은 각각 고유한 검증 방식을 가지며 실제 작업의 상호 보완적인 측면을 탐구합니다. 이 모든 것은 통일된 태스크 디렉터리 형식으로 패키징되어, 통일되고 재현 가능한 프로토콜 하에 두 가지 에이전트 도구(CodeBuddy Code 및 Claude Code)에서 실행됩니다. 전체 공개를 통해 벤치마크는 종단 간 재현 가능하고 직접 감사 가능해집니다. 이는 제3자가 각 태스크를 다시 실행하고 내용을 검사할 수 있기 때문입니다. 각 하위 집합이 서로 다른 채점 도구를 사용하므로 점수는 하위 집합 간에 비교 불가능하며, 제품군 전체 평균은 보고되지 않습니다. 여러 모델 계열에 걸친 교차 모델 리더보드를 보고합니다.
English
We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation framework for constructing and running distribution-informed coding-agent tasks across four work domains - Code, Web, Office, and Security. Rather than adapting public issue text, every task is reverse-engineered from a real commit, pull request, or business scenario and rewritten as a short, colloquial, role-played request, so that a task's prompt is not recoverable by web-searching the underlying issue, pull request, or commit thread. Because the dataset is released openly - task directories, environment images, evaluation harness, tests, and reference solutions - contamination resistance rests on this construction together with dataset versioning rather than on secrecy. The four subsets - repository-level engineering, front-end development, office and business workflows, and red-/blue-team security - probe complementary facets of real work, each with its own verification style. All are packaged in a uniform task-directory format and run, under a uniform and reproducible protocol, on two agent harnesses (CodeBuddy Code and Claude Code); the full open release makes the benchmark reproducible end to end and directly auditable, since any third party can re-run each task and inspect its content. Because each subset uses a different scoring instrument, scores are not comparable across subsets and the suite reports no suite-wide average. We report a cross-model leaderboard across several model families.