Skip to content
Longterm Wiki

Beth Barnes - METR

web

Credibility Rating

4/5
High(4)

High quality. Established institution or organization with editorial oversight and accountability.

Rating inherited from publication venue: METR

Metadata

1 FactBase fact citing this source

Cached Content Preview

HTTP 200Fetched Oct 3, 20267 KB
Beth Barnes - METR 

 

 

 
 

 

 

 
 
 
 

 

 

 
 
 
 
 
 
 

 

 

 
 

 

 

 

 
 
 
 
 
 
 
 
 
 
 
 

 
 
 
 
 
 
 
 
 
 
 
 
 Our Work 
 
 
 
 
 
 
 
 
 
 
 Research 
 

 
 
 
 
 
 
 
 Notes 
 

 
 
 
 
 
 
 
 Updates 
 

 
 
 
 
 
 
 
 Risk Assessment 
 

 
 
 
 

 
 
 
 
 
 
 
 
 About 
 

 
 
 
 
 
 
 
 
 Donate 
 

 
 
 
 
 
 
 
 
 Careers 
 

 
 
 

 
 
 Search 
 

 
 

 
 
 
 

 

 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 -->
 
 
 
 

 

 
 
 
 
 
 
 
 
 
 
 
 
 
 
 Our Work 
 
 

 
 
 
 
 
 
 
 
 
 

 
 Research 
 
 

 

 
 
 
 
 
 
 
 
 

 
 Notes 
 
 

 

 
 
 
 
 
 
 
 
 

 
 Updates 
 
 

 

 
 
 
 
 
 
 
 
 

 
 Risk Assessment 
 
 

 

 
 

 

 
 
 
 
 
 
 
 
 
 
 About 
 

 
 
 
 
 
 
 
 
 
 
 Donate 
 

 
 
 
 
 
 
 
 
 
 
 Careers 
 

 
 

 

 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

 
 

 

 
 

 
 Menu 
 

 
 

 

 
 
 
 
 
 × 
 
 

 
 Beth Barnes 
 
 
 
 
 

 
 
 
 
 
 

 
 
 
 Beth Barnes

 
 Founder, CEO
 
 
 
 
 Beth Barnes is the Founder and CEO of METR, where she oversees a growing technical team that designs and carries out evaluations of generative AI models. Beth previously worked with DeepMind’s Chief Scientist on scaling laws for forecasting deep learning progress and at OpenAI where she helped OpenAI develop safety targets and evaluated scalable oversight techniques for alignment and code models for misalignment before release. She has written and presented on a range of technical and theoretical issues related to aligning machine learning systems with human values.

 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

 
 By Beth Barnes

 
 
 
 
 
 

 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 Implementing and Evaluating a Basic Per-Action Monitor for Safer Evals
 
 

 
 September 27, 2026
 

 
 
 Reilly Haskins, Rif A. Saurous, Nate Rush, Neev Parikh, and Beth Barnes lay out the claims that would need to hold for their per-action blocking monitor to be effective, the evidence for each claim, and the gaps that remain.

 
 
 
 
 Read more 
 
 
 
 
 
 

 
 
 
 
 
 

 
 
 
 
 
 
 
 
 
 
 
 
 
 
 Frontier Risk Report (February to March 2026)
 
 

 
 May 19, 2026
 

 
 
 A pilot assessment of rogue deployment risk at frontier AI companies. Starting in February 2026, METR conducted a pilot exercise to assess misalignment risks from AI agents used inside frontier AI developers, with participation from Anthropic, Google, Meta, and OpenAI.

 
 
 
 
 Read more 
 
 
 
 
 
 

 
 
 
 
 
 

 
 
 
 
 
 
 
 
 Review of the "Risks from automated R&D" section in the Anthropic Risk Report (February 2026)
 
 

 
 May 8, 2026
 

 
 
 External review from METR of the "Risks from automated R&D" section in Anthropic's February 2026 Risk Report

 
 
 
 
 Read more 
 
 
 
 
 
 

 
 
 
 
 
 

 
 
 
 
 
 
 
 
 Review of the Anthropic Sabotage Risk Report: Claude Opus 4.6
 
 

 
 March 12, 2026
 

 
 
 External review from METR of Anthropic's Sabotage Risk Report for Claude Opus 4.6

 
 
 
 
 Read more 
 
 
 
 
 
 

 
 
 
 
 
 

 

... (truncated, 7 KB total)
Resource ID: 65513014144f0634 | Stable ID: sid_qskIvjdLCw