Pioneers Insight Method Research Author
Back to Pioneers
Buck Shlegeris
Developers 1 Curated Dialogues

Buck Shlegeris

Key Views & Dialogues

AI Control: Using Untrusted Systems Safely with Buck Shlegeris, Redwood Research (80,000 Hours Pod)

  • 🗓️ Date2025-05-04 | 🎙️ Show:The Cognitive Revolution

AI control offers a deployment thesis for extracting useful work from potentially misaligned systems without first making them trustworthy. Cheap monitors, suspicious-action replacement, resampling, and escalating audits can improve safety within roughly 5% budget or productivity constraints, though stronger scaffolds require organizational buy-in. The binding risk is execution: frontier labs may have a manageable list of controls but sharply lower appetite to implement them as capabilities accelerate.

View Dialogue Notes & Key Takeaways
  • AI control is a practical deployment thesis for models that may already be egregiously misaligned, not an attempt to make them trustworthy first. Buck Shlegeris’s target scenario is an AI lab with roughly 1,000 human researchers and 100,000 model copies operating around 16× human speed, with access to code, experiments, model weights, and compute. The core asset at risk is the data center—“that’s where the compute is”—and the objective is to extract useful work while making hacking, weight exfiltration, research sabotage, and data poisoning unlikely to succeed.

  • The most deployable defense is an escalation stack: cheap models inspect every action, suspicious actions are replaced or resampled, and only the extreme tail reaches expensive audits and humans. If humans can inspect only one action in 1,000, the monitor must operate near a 0.1% false-positive rate, where reliably detecting every attack is unrealistic. Blocking a much wider set of merely suspicious actions lets defenders “live to fight another day,” forcing the model to attack repeatedly and creating more chances to catch it.

  • AI’s numerical advantage is offset by a security property humans do not offer: developers can largely control what information enters and leaves each instance. Unlike corporate spies, who can coordinate outside work, models may not know the date, whether they are in production, or whether a prior action has been replayed. That makes mass honeypotting, trajectory resampling, simulated escape attempts, and repeated sting operations possible—although correlated behavior means that another copy of the same model cannot be treated as an independent reviewer.

  • The commercial adoption constraint is likely to be cost and performance, not theoretical elegance. Buck imagines safety teams receiving only 5% of budget or being allowed no more than a 5% productivity loss; even replacing the sketchiest 2% of actions might deliver substantial safety. Earlier models are usually far cheaper to run, so always-on monitoring can be economical, but more powerful techniques such as paraphrasing, context restriction, and “forced factored cognition” require deeper changes to agent scaffolds and stronger organizational buy-in.

  • Catching one escape attempt is valuable even if it does not produce a global pause. A lab can let the model believe it succeeded, observe its reserved exploits, generate positive examples for internal-state probes, and feed each new attempt back into detection systems; Buck argues it is “seriously plausible” the lab could feel safer on Wednesday than it did on Monday after catching an attempt Tuesday. Yet proof may not persuade competitors or regulators: a prior of 30% misalignment only moves 3× when updated to certainty, while skeptics could call the incident an aberration or even allege the lab deliberately trained the behavior.

  • The largest residual risks are opaque reasoning, long-context coordination, and chronic sabotage that ordinary control metrics do not capture. Transparent chains of thought would make scheming much easier to spot—Buck says if AI systems automating AGI development still require long, readable reasoning, “we are probably just totally fine”—and current tests found o1 not scarily capable of scheming even with chain of thought. But fast capability gains, 10× more agents, hidden recurrence, covert inter-model codes, and selectively poor work in domains where evaluators cannot distinguish good from bad could erode that advantage rapidly.

  • Execution inside frontier labs is now the binding variable in Buck’s worldview. Five years ago he expected “galaxy brain fundamental insights”; now he sees “a list of 40 things” that are individually manageable, while having updated sharply downward on companies’ appetite to implement them. That creates unusually high leverage for perhaps 10 well-placed insiders, exportable outside research, common standards, and a prepared portfolio of roughly 50 projects that more capable—but controlled—AI researchers could execute later.

  • 🔗 Original source & video: AI Control: Using Untrusted Systems Safely with Buck Shlegeris, Redwood Research (80,000 Hours Pod)

Listen to full conversation →