BenchFlow builds the environments AI agents learn in.
Four products stopped looking like four companies. SkillsBench, ClawsBench, PostTrain Arena and the runtime now share one ground, one type system and one URL space — and the teal is gone.
Everything becomes a route under BenchFlow
SkillsBench
Does loading a skill make an agent better at real work? 86 tasks, 11 domains, 25 model × harness configs.
ClawsBench
Five mock workplaces wire-compatible with the real Workspace and Slack APIs. Capability and safety, measured separately.
PostTrain Arena
You contribute environments; we post-train on them and score what generalizes across eight under-served domains.
BenchFlow runtime
One Scene-based lifecycle for single-agent, multi-agent and multi-round evals. TaskMiner and the capability viewer sit on top of it.
One family, three jobs
tasks 86 · domains 11 · trials 7,224
Data is the bottleneck. Environments are the new data.
Labels
Image tags, span annotations, yes/no labels.
Post-training
SFT, preferences, reward labels, short trajectories.
Environments
Stateful workplaces with services, files, tools, verifiers, replay.