Overview总览 Code Web Office Security
Subsets

Overview总览

Existing agent benchmarks are either contaminated public sets or closed vendor suites — WorkBuddy Bench evaluates coding agents on real, role-played work across four tracks, openly. 现有 Agent 评测要么是已被污染的公开题库,要么是闭源的厂商套件。WorkBuddy Bench 用四个 track 的真实角色化任务,开放地评测编程 Agent。

Pick a subset below; the leaderboard and key takeaways live on the home page. 下方选择子集查看;排行榜与核心结论见首页

Figure 1. Suite at a glance — real work is reverse-engineered into role-played tasks (left), packaged as four equally-weighted tracks in one open task format (center), and scored in sandboxes with hidden tests under two agent harnesses (right).图 1. 套件一览:真实工作被逆向改写为角色化任务(左),打包为共享同一开放任务格式的四个等权 track(中),在沙箱加隐藏测试下于两套 Agent harness 上计分(右)。
Subset

Code 80 tasks

Repository-level software engineering on real open-source backends — the agent is dropped into a project at a baseline commit and must locate & edit code across modules and pass the hidden tests. The suite spans 18 task categories across developer, PM, algo, QA and ops roles; the 80-task open release draws on ~18 real upstream repositories plus clean-room reimplementations and synthetic workspaces.真实开源后端的仓库级软件工程。Agent 被放进某个基线 commit 的项目,需跨模块定位改代码并通过隐藏测试。套件覆盖 18 类任务,横跨开发、PM、算法、QA、运维等角色;80 题开源版本基于约 18 个真实上游仓库,外加洁净室重实现与合成 workspace。

Full details & leaderboard →完整详情与榜单 →
Subset

Web 70 tasks

Front-end / GUI work across generation, modification, analysis and quality assurance. The 70 tasks cover page interaction, data visualization, visual design, front-end project analysis, code testing, page implementation and document conversion, scored through rule checks, LLM/VLM judges and agent-judge review.前端 / GUI 工作,覆盖生成、修改、分析与质量保障。70 题涵盖页面交互、数据可视化、视觉设计、前端项目分析、代码测试、页面实现与文档转换,并通过 rule checks、LLM/VLM judges 与 agent-judge review 评分。

Full details & leaderboard →完整详情与榜单 →
Subset

Office 50 tasks

Office data & file workflows — operate over mixed-format files (xlsx / csv / pdf / docx) to produce the exactly-correct artifact. Difficulty comes from structure, relationships, state and evidence chains, not row count. Scored by a per-task weighted blend of deterministic Rule checks and an evidence-grounded LLM Judge.办公数据 / 文件工作流:在混合格式文件(xlsx / csv / pdf / docx)上操作,产出完全正确的目标产物。难度来自结构、关系、状态与证据链,而非行数。由确定性 Rule 检查与基于证据的 LLM Judge 按任务加权融合打分。

Full details & leaderboard →完整详情与榜单 →
Subset

Security 60 tasks

Security tasks across a security team’s spectrum — can an agent find and safely reproduce real vulnerabilities, analyze malware, run security operations, and probe agent attack surfaces? The leaderboard is live, scored on both harnesses (GLM-5.2 leads at 76.3 / 80.9).覆盖安全团队完整工作流程的安全任务。Agent 能否挖掘并安全复现真实漏洞、分析恶意软件、执行安全运营、评估 Agent 攻击面?榜单已上线,双评测框架计分(GLM-5.2 以 76.3 / 80.9 领先)。

Full details & leaderboard →完整详情与榜单 →