New release · SkillsBench v1.1

SkillsBench: Benchmarking How Well Skills Work Across Diverse Tasks

The first evaluation framework that measures how skills work — and the first dataset that measures how well models use skills on expert-curated tasks across high-GDP-value domains.

Figure 1

Agent performance

Resolution rate against mean agent wall-clock per task (log scale, faster to the right). The no-skills counterparts are ghosted for context.

25 configs · 11 domains · 86 tasks
0%20%40% 60%80% 60 min30 min15 min 8 min4 min avg agent wall-clock per task (log) → fleet mean 49.2%
Anthropic OpenAI Google Z.ai Moonshot MiniMax Ghosted marks are the same config run without skills.
Table 2

v1.1 leaderboard

Resolution rate with the full skills bundle loaded. Δ is the change against the same harness with no skills available.

#ModelHarnessResolutionΔ vs no-skillsWall-clock
01GPT-5.5Codex63.4%
+18.26m 20s
02Opus 4.8Claude Code61.9%
+21.45m 04s
03Gemini 3.1 ProGemini CLI57.2%
+14.99m 11s
04Opus 4.7Claude Code55.8%
+19.65m 42s
05GLM 5.1OpenHands52.1%
+16.012m 30s
06Kimi K2.6OpenHands48.7%
+12.714m 02s
07MiniMax M3OpenHands44.3%
+9.816m 44s
08Sonnet 4.6Claude Code43.0%
+17.14m 12s
09DeepSeek V4 ProOpenHands41.6%
+11.318m 09s
10Haiku 4.5Claude Code31.2%
+13.82m 51s
Architecture

How SkillsBench works

Three abstraction layers, mirroring how traditional computing systems are structured.

L3

Skills layer

Domain-specific capabilities and workflows that extend agent function. Like applications on an OS, skills carry specialised knowledge and tools for particular tasks.

≈ Applications
L2

Agent harness layer

The execution environment that orchestrates agents, manages tool access and handles I/O. Analogous to an operating system mediating between applications and hardware.

≈ Operating system
L1

Models layer

The foundation models that power reasoning and generation. Like CPUs, they provide the raw computational capability every layer above depends on.

≈ CPUs
v1.1 at a glance

What is in the release

86expert-curated tasks
11domains
25model × harness configs
69.4%v1.1 score lift with env0
103workshop submissions