Automating AI Safety Research
We build benchmarks, datasets, and infrastructure that help speed up AI safety research.
Our Mission
We are a research group focused on building open infrastructure for AI safety research. Our work is motivated by the belief that AI systems will become increasingly capable and autonomous, and that we need to develop the infrastructure, now, to keep pace with capabilities
What We Build
Benchmarks
We build benchmarks to evaluate how AI agents can contribute to AI safety research.
Datasets
1.1M enriched papers. 129K research repositories. 778K code functions. The raw material for studying how AI systems interact with real scientific work.
Infrastructure
Runtimes for structured agent workloads. Orchestration topologies for recursive improvement loops. The scaffolding to run experiments at scale.
Research
ARIA Benchmark: How Much Machine Learning Do AI Models Actually Know?
A suite of five closed-book benchmarks probing the ML knowledge that frontier language models have internalized during training.
ArXiv Research Code Dataset: 129K Research Repositories
A collection of 4.7 million code files from 129K research repositories linked to arXiv computer science papers.
ArXivDLInstruct: 778K Research Code Functions for Instruction Tuning
A dataset of 778,152 functions extracted from arXiv-linked research code, each paired with instruction prompts, for training...
DeltaML-Bench: Evaluating Machine Learning Agents on Real-World Research Repositories
A 48-task benchmark for evaluating autonomous ML experimentation in real research repositories.
ML Research Benchmark: Can AI Agents Do Real ML Research?
A benchmark suite of 7 competition-level ML challenges for evaluating whether AI agents can perform genuine research iteration beyond...
About
Algorithmic Research Group builds tools and infrastructure for AI safety research. Benchmarks for evaluating autonomous agents. Datasets for studying how models fail. Runtimes for running agent workloads at scale.
