BenchFlow builds the environments AI agents learn in.
A frontier environment lab for AI agents. We ship SkillsBench, ClawsBench, and the BenchFlow runtime.
What we ship
All research →SkillsBench
The first benchmark for whether procedural skills — instructions, scripts and references an agent loads on demand — make agents better at real work. 86 tasks, 11 domains.
ClawsBench
Five mock workplaces — Gmail, Calendar, Drive, Docs, Slack — wire-compatible with the upstream gws and Slack APIs. Production agents and skills run unchanged against a safety-evaluable replica.
BenchFlow runtime
The agent simulation runtime. One Scene-based lifecycle for single-agent, multi-agent and multi-round evals. Sandboxed, hardened against reward hacking, full trajectory capture.
Data is the bottleneck. Environments are the new data.
AI data went from labels to post-training trajectories to environments. Models in 2026 do not get better from more static prompts — they get better from running through realistic environments and being judged on the whole workflow.
Labels
Image tags, span annotations, yes/no labels.
Post-training
SFT, preferences, reward labels, short trajectories.
Environments
Stateful workplaces with services, files, tools, verifiers, replay.
Ecosystem
2026Agent Skills ’26 workshop
First workshop on agent skills. Speakers: Dawn Song, Ross Taylor, Kanav Garg (DeepMind), Yu Su. Live SkillsBench design challenge.
agentskills-workshop.org ↗SkillsBench 1.0 launch party
Presented by Google DeepMind — the afterparty to the CAIS workshop. We announced SkillsBench 1.0: 100+ expert-curated tasks, built with Kaggle.
skillsbench.ai ↗