Automating AI Safety Research

We build benchmarks, datasets, and infrastructure that help speed up AI safety research.

We are a research group focused on building open infrastructure for AI safety research. Our work is motivated by the belief that AI systems will become increasingly capable and autonomous, and that we need to develop the infrastructure, now, to keep pace with capabilities


Benchmarks

We build benchmarks to evaluate how AI agents can contribute to AI safety research.

Datasets

1.1M enriched papers. 129K research repositories. 778K code functions. The raw material for studying how AI systems interact with real scientific work.

Infrastructure

Runtimes for structured agent workloads. Orchestration topologies for recursive improvement loops. The scaffolding to run experiments at scale.

Algorithmic Research Group builds tools and infrastructure for AI safety research. Benchmarks for evaluating autonomous agents. Datasets for studying how models fail. Runtimes for running agent workloads at scale.