SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Published:

Authors: Xiangyi Li, Yimin Liu, Wenbo Chen, et al., including Zelin Tan.

SkillsBench evaluates whether structured Agent Skills improve LLM agents on expertise-heavy tasks. Its current release contains 87 tasks across eight domains and finds that curated Skills substantially improve average pass rates across model and agent-harness configurations.

Links: Paper · Website · Code