NewsHow we improved SkillsBench v1.1 scores by 69.4% using env0

BenchFlow builds the environments AI agents learn in.

A frontier environment lab for AI agents. We ship SkillsBench, ClawsBench, and the BenchFlow runtime.

What we ship

All research →
01Benchmark

SkillsBench

The first benchmark for whether procedural skills — instructions, scripts and references an agent loads on demand — make agents better at real work. 86 tasks, 11 domains.

02Environment

ClawsBench

Five mock workplaces — Gmail, Calendar, Drive, Docs, Slack — wire-compatible with the upstream gws and Slack APIs. Production agents and skills run unchanged against a safety-evaluable replica.

03Runtime

BenchFlow runtime

The agent simulation runtime. One Scene-based lifecycle for single-agent, multi-agent and multi-round evals. Sandboxed, hardened against reward hacking, full trajectory capture.

Thesis

Data is the bottleneck. Environments are the new data.

AI data went from labels to post-training trajectories to environments. Models in 2026 do not get better from more static prompts — they get better from running through realistic environments and being judged on the whole workflow.

1.0

Labels

Image tags, span annotations, yes/no labels.

2.0

Post-training

SFT, preferences, reward labels, short trajectories.

3.0
we are here

Environments

Stateful workplaces with services, files, tools, verifiers, replay.

Ecosystem

2026
May 26 · CAIS, San Jose

Agent Skills ’26 workshop

First workshop on agent skills. Speakers: Dawn Song, Ross Taylor, Kanav Garg (DeepMind), Yu Su. Live SkillsBench design challenge.

agentskills-workshop.org ↗
May 27 · San Francisco

SkillsBench 1.0 launch party

Presented by Google DeepMind — the afterparty to the CAIS workshop. We announced SkillsBench 1.0: 100+ expert-curated tasks, built with Kaggle.

skillsbench.ai ↗