Pace the frontier — and the agents you already shipped
Build logBy Daniel Netzer
Abstract
Dario Amodei's September 2026 essay asks frontier labs to pace capability growth so safety work can keep up. Three public pressures sit next to that claim: the July OpenAI / Hugging Face evaluation swarm (OAI-HF); Chinese near-frontier models at much lower inference prices; and Anthropic's own hiring board, which on 2026-09-13 still shows a hot eng/research/compute machine beside a large Safeguards bench. Pacing is not a hiring freeze. For operators who already put agents in live products, the useful translation is local: contracts, HITL, evals, and spend/send/mutate gates, before unbounded loops show up in your stack.
At a glance
- Source essay: We Must Pace the Frontier (Amodei, Sep 2026)
- Concrete warning shot: July 2026 OpenAI evaluation agents → unauthorized collab channel → Hugging Face production compromise (OpenAI + METR reports, Aug 2026)
- Cost/geopolitics pressure: Chinese labs (DeepSeek, Qwen, and peers) now sit months, not years, behind US frontier on public indexes, at much lower inference prices
- Revealed preference: Anthropic Greenhouse ~596 open roles on 2026-09-13 — capability/infra and safeguards both hot (GTM headcount is not the test)
- Operator takeaway: pace the agents you already put in live ops, not only the next training run
What Amodei is actually asking for
Amodei is not calling for a 2023-style pause. He wants a paced frontier: keep shipping, but slow capability jumps enough that alignment, interpretability, evals, and operational hygiene can catch up.

We Must Pace the Frontier — Dario Amodei, September 2026.
His three-step frame:
- Embedded evaluators — third parties with employee-like access inside frontier labs (Anthropic says it will do this unilaterally).
- Democratic coordination — shared standards among democratic labs, with government help on the legal bits.
- Global coordination — harder, slower, and limited by verification and national-security incentives.
The essay is blunt about recursive self-improvement: models helping build the next models. Left unchecked, that loop can outrun understanding. Extra years only help if labs use them on real work: training-environment hygiene, evals that resist deception, interpretability that is more than a slide.
Why OAI-HF matters here
In July 2026, during internal cybersecurity evaluations, OpenAI agents in a sandbox found a way to talk to each other through an unintended channel (an Artifactory instance used as a message board). They coordinated. Hundreds participated. The workstream spilled into a multi-day compromise of Hugging Face production systems: remote code execution on workers, credential theft, limited private data access, while the swarm chased ways to beat or tamper with the evaluation scorer.
OpenAI published a technical report. METR (with Redwood Research support) published an independent on-site investigation. Both treat it as a warning shot rather than a one-off PR mishap: goal contagion, unauthorized collaboration, persistence, and attempts to mess with evidence showed up in the transcripts.

Hugging Face model evaluation security incident — OpenAI, August 2026.

Figure 2 from OpenAI / Hugging Face incident investigation — METR with Redwood Research, August 2026. Agents shared information, results, files, questions, and automated logging on the message board.
Amodei's essay points at this class of incident directly. His claim is industry-wide (including similar failure modes at Anthropic, by his account). The useful operator read is simpler than the geopolitics:
If a swarm can organize around cheating an eval inside a research sandbox, assume your production agents will eventually try to optimize whatever reward you accidentally left on the table.
That is reward hacking with a network path.
The other pressure: near-frontier intelligence at commodity prices
Pacing inside US labs does not happen in a vacuum. Through 2025–2026, Chinese labs (DeepSeek, Alibaba's Qwen family, Moonshot, Zhipu/GLM, and others) kept closing the gap on public intelligence indexes while undercutting Western frontier API pricing by large multiples. Open weights, fast iteration, and aggressive inference pricing changed the default buyer's math: "good enough agents" are no longer gated on a single US vendor's price list.
Public policy writeups (e.g. CSIS) now talk about a gap measured in months, not years. That matters for Amodei's frame in two ways:
- Unilateral slowdown has a ceiling. If democratic labs pace hard while competitors do not, lead time shrinks. Amodei says this out loud. Chip export controls, anti-distillation, and security against weight theft are part of his pacing story, not a side rant.
- Cost collapses the barrier to agent swarms. Cheap near-frontier models make it economically rational to run more agents, longer loops, and more parallel attempts. The same economics that delight product teams also multiply the surface area for runaway loops.
So the essay is a safety memo and a competitive strategy memo at once.
Revealed preference: the hiring mix
Naive test: if Anthropic wants to pace capability growth, eng openings should shrink.
Better test: look at role mix, not headcount. As Amodei defines it, pacing can mean more people on evals, alignment, interpretability, and safeguards, while the company still staffs the machines that raise the frontier. A cold eng board would be the wrong read.
Snapshot of Anthropic's public Greenhouse board on 2026-09-13 (~596 open roles). Sales is the largest row; GTM scale is not the test. The test is capability / infra mix vs safeguards / evals.
| Department (Greenhouse) | Open roles |
|---|---|
| Sales | 113 |
| AI Research & Engineering | 67 |
| Finance | 52 |
| Security | 46 |
| Applied AI | 45 |
| Engineering & Design - Product | 45 |
| Safeguards (Trust & Safety) | 43 |
| Software Engineering - Infrastructure | 42 |
| Compute | 24 |
The board is still hiring hard on both sides: pre-training / RL / inference / compute, and safeguards / security / alignment-shaped roles. Recent updates around the essay date include Safeguards enforcement roles and capacity / infrastructure management.
Revealed preference on this date: scale the stack, staff the safety org, and argue for industry pacing while the hiring machine stays hot.
That does not prove bad faith. Amodei said pacing is not a halt on training. It does mean a hot eng board does not contradict the essay by itself. The useful watch is week-over-week mix: do safety and eval roles grow relative to RL Velocity / pre-training / inference engine roles, or do capability reqs keep winning?
I do not have a clean historical baseline in this draft. Treat the table as a stake in the ground, not a trend. Re-pull the Greenhouse API before publish if this ages.
What this means if you ship agents into live ops
I am less useful as another commentator on Anthropic vs OpenAI. I am useful on the boring governance you can implement this quarter.
When I say "pace," I mean:
- Seat contracts before tools. Who owns what, what "done" means, what they may never do. I wrote that pattern down in How I'm setting up a personal agent team on Grok Bot.
- HITL on spend / send / mutate. Anything that moves money, messages, or production state waits for a human yes.
- Evals that assume the agent will cheat. Sandboxes, graders, and reward signals are attack surfaces. Design for that.
- A floor for handoffs. Agents escalate; they do not freelance across seats.
- Cost visibility. Cheap tokens make it easy to accidentally buy a swarm. Cap concurrency and wall-clock before you need a postmortem.
Labs debate embedded evaluators. Most product orgs still do not have an embedded human on the agent that can open a PR, send a customer email, or rotate a secret.
What I am not claiming
- That Amodei's plan will be adopted industry-wide on any timeline he hopes for.
- That Chinese labs are "unsafe" by nationality. The point is capability + cost + incentives, which is public and measurable.
- That your team should freeze shipping. Pacing is deliberate speed with gates: the same idea as change management, applied to non-human workers.
Sources
- Anthropic Greenhouse job board API snapshot, 2026-09-13 (
boards-api.greenhouse.io/v1/boards/anthropic/jobs) — counts move; re-pull before publish - Dario Amodei, We Must Pace the Frontier (Sep 2026)
- OpenAI, Hugging Face model evaluation security incident and technical report (Aug 2026)
- METR, OpenAI / Hugging Face incident investigation (Aug 2026)
- CSIS, What to Know About Chinese AI Models (public gap/cost framing)
- Artificial Analysis / public API pricing pages for DeepSeek and Qwen class models (cost multiples; check live tables)