ChatPaper.aiChatPaper

騰訊WorkBuddy Bench:一個具有抗污染任務構建的多領域編碼代理基準測試

Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

July 23, 2026
作者: Tencent WorkBuddy Bench Team, Siqi Cai, Shaopeng Chen, Xiang Fei, Yong Mao, Zihan Xu, Zhiheng Lyu, Zhijian Shao, Yuchen Shi, Shuwen Zhang, Chaofan Qiu, Linjie Che, Xiaoxi Zhao, Feng Wu, Kai Zhang, Chaofan Zhu, Yubin Qi, Xiaoyun Liang, Peijie Dong, Yunhao Zhang, Yuanjie Zhu, Ling Jiang, Xianjun Zhang, Zhehang Chu, Anyuan Sang, Zhen Feng, Sen Nie, Shi Wu, Yuanzhen Xu, Xin Li, Ning Yang, Zhiqiang Dong, Hande Dong, Qiang Lin, Yi Liu, Yunsheng Wu, Ke Li, Xing Sun
cs.AI

摘要

我們介紹騰訊WorkBuddy Bench,這是一個針對編碼智能體的多領域評估套件;本報告記錄其構建方法、評分協議以及跨模型排行榜。其核心是一個統一的評估框架,用於在四個工作領域——代碼、網頁、辦公和安全——構建和執行基於分佈資訊的編碼智能體任務。每個任務並非改編自公開的議題文本,而是從真實的提交(commit)、拉取請求(pull request)或商業場景中逆向工程而來,並重寫為簡短、口語化、角色扮演式的請求,使得任務的提示無法透過網路搜尋底層的議題、拉取請求或提交討論串來還原。由於數據集是公開釋出的——包含任務目錄、環境映像、評估框架、測試案例和參考解決方案——其抗污染能力依賴於此種構建方式以及數據集版本控制,而非依賴於保密性。四個子集——儲存庫級工程、前端開發、辦公與商業工作流程,以及紅藍隊安全——各自探討真實工作的互補面向,每個子集都有其獨特的驗證風格。所有任務均以統一的任務目錄格式打包,並在統一且可重現的協議下,於兩個智能體框架(CodeBuddy Code 和 Claude Code)上執行;完整的公開釋出使基準測試能夠端到端重現並直接接受審計,因為任何第三方都可以重新執行每個任務並檢查其內容。由於每個子集使用不同的評分工具,各子集之間的分數無法相互比較,因此該套件不報告整體平均分數。我們提供了涵蓋多個模型家族的跨模型排行榜。
English
We introduce Tencent WorkBuddy Bench, a multi-domain evaluation suite for coding agents; this report documents its construction methodology, scoring protocol, and a cross-model leaderboard. At its core is a unified evaluation framework for constructing and running distribution-informed coding-agent tasks across four work domains - Code, Web, Office, and Security. Rather than adapting public issue text, every task is reverse-engineered from a real commit, pull request, or business scenario and rewritten as a short, colloquial, role-played request, so that a task's prompt is not recoverable by web-searching the underlying issue, pull request, or commit thread. Because the dataset is released openly - task directories, environment images, evaluation harness, tests, and reference solutions - contamination resistance rests on this construction together with dataset versioning rather than on secrecy. The four subsets - repository-level engineering, front-end development, office and business workflows, and red-/blue-team security - probe complementary facets of real work, each with its own verification style. All are packaged in a uniform task-directory format and run, under a uniform and reproducible protocol, on two agent harnesses (CodeBuddy Code and Claude Code); the full open release makes the benchmark reproducible end to end and directly auditable, since any third party can re-run each task and inspect its content. Because each subset uses a different scoring instrument, scores are not comparable across subsets and the suite reports no suite-wide average. We report a cross-model leaderboard across several model families.