logo
Tencent WorkBuddy Bench
Grounded in real usage · contamination‑resistant · open
Technical Report · Fig. 1
1Construction— from real work to a task
Real commit
Pull request
Business scenario
Security case
Sourced from real repositories, tickets & vulnerability reportsroot cause & original fix withheld from the agent
Grounded in real‑usage analysis — task mix matched to production scenario & query‑intent distributions; no raw prompts, sessions or logs shipped
reverse‑engineered into colloquial, role‑played asks — every subset
developer
Developer · Code“Don’t let cleanup leak exceptions at exit.”
Security researcher · Security“Locate the flaw, then reproduce it safely with a verified PoC.”
pm
PM · Office“Reconcile these sheets and flag the mismatches.”
2The Benchmark— four real‑work tracks
One suite, four independent tracks
tasks/<id>/ instruction.md task.toml environment/ tests/
Code
80
Repo‑scale
modify & fix inside real repos
Web
70
Interactive
build & verify interactive UI
Office
50
Workflow
spreadsheets, docs & data flows
Security
60
Red + blue
Red‑team38 attack
Blue‑team22 defense
find, reproduce & analyze
3Evaluation— dual harness
Primary
CodeBuddy Code
Alternate
Claude Code
Isolated sandbox
Hidden tests
Numeric reward
0–1
Codehidden‑test reward Webweighted rule / VLM / judge Officerule + LLM‑judge blend Securityprogrammatic scoring
Current leader · per track
CodeClaude Opus 4.8
74.4
WebClaude Opus 4.8
68.1
OfficeClaude Opus 4.8
82.4
SecurityGLM‑5.2
76.3
Top score per track, CodeBuddy Code harness. Tracks use different metrics — scores are not comparable across tracks.