SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks

Published in NeurIPS 2026, 2026

Authors: Xiangyi Li, Yimin Liu, Wenbo Chen, et al., including Zelin Tan.

SkillsBench evaluates whether structured Agent Skills improve LLM agents on expertise-heavy tasks. Its current release contains 87 tasks across eight domains and finds that curated Skills substantially improve average pass rates across model and agent-harness configurations.

Links: Paper · Website · Code · NeurIPS 2026