SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks
Published:
Authors: Xiangyi Li, Yimin Liu, Wenbo Chen, et al., including Zelin Tan.
SkillsBench evaluates whether structured Agent Skills improve LLM agents on expertise-heavy tasks. Its current release contains 87 tasks across eight domains and finds that curated Skills substantially improve average pass rates across model and agent-harness configurations.
