Skip to content
Longterm Wiki

Alignment Research Center - Wikipedia

reference

Credibility Rating

3/5
Good(3)

Good quality. Reputable source with community review or editorial standards, but less rigorous than peer-reviewed venues.

Rating inherited from publication venue: Wikipedia

ARC is a key organization in the AI safety landscape; its ARC Evals team conducts pre-deployment capability evaluations for frontier AI labs including Anthropic and OpenAI, making it directly relevant to deployment safety and governance discussions.

Metadata

Importance: 55/100wiki pagereference

Summary

Wikipedia overview of the Alignment Research Center (ARC), a nonprofit AI safety research organization founded in April 2021 by Paul Christiano. ARC focuses on developing scalable alignment methods, evaluating dangerous AI capabilities, and ensuring advanced AI systems are safe and beneficial. It has expanded from theoretical work into empirical research, industry collaborations, and policy.

Key Points

  • •Founded in April 2021 by former OpenAI researcher Paul Christiano, based in Berkeley, California.
  • •Develops scalable methods for training AI systems to behave honestly and helpfully, analyzing how alignment techniques could break down.
  • •ARC Evals (started by Beth Barnes) focuses on evaluating capabilities and alignment of advanced AI models, including red-teaming frontier models.
  • •Funded primarily by Open Philanthropy; notably returned a $1.25M FTX Foundation grant after the FTX collapse on ethical grounds.
  • •Expanding scope from theoretical alignment research into empirical work, industry partnerships, and AI policy engagement.

Cited by 2 pages

Cached Content Preview

HTTP 200Fetched Oct 3, 20268 KB
From Wikipedia, the free encyclopedia 
 
 
 
 
 AI safety research organization 
 Not to be confused with Arc Institute . 
 "},"fields":{"wt":"[[AI alignment]] and [[AI safety|safety research]]"},"location":{"wt":"[[Berkeley, California]]"},"website":{"wt":"{{url|https://www.alignment.org/|alignment.org}}"}},"i":0}}]}'> Alignment Research Center Formation April   2021 ; 5   years ago   ( 2021-04 ) Founder Paul Christiano Type Nonprofit research institute Legal   status 501(c)(3) tax exempt charity Location Berkeley, California Fields AI alignment and safety research Website alignment.org 

 The Alignment Research Center ( ARC ) is a nonprofit research institute based in Berkeley, California , dedicated to the alignment of advanced artificial intelligence with human values and priorities. [ 1 ] Founded by former OpenAI researcher Paul Christiano , ARC established an evaluation team to study the potentially harmful capabilities of present-day AI models. [ 2 ] [ 3 ] This team became the independent nonprofit METR in December 2023. [ 4 ] 

 History and research

 [ edit ] 
 ARC's mission is to ensure that powerful machine learning systems of the future are designed and developed safely and for the benefit of humanity. It was founded in April 2021 by Paul Christiano and other researchers focused on the theoretical challenges of AI alignment. [ 5 ] ARC aims to develop scalable methods for training AI systems to behave honestly and helpfully. A key part of its methodology is considering how proposed alignment techniques might break down or be circumvented as systems become more advanced. [ 6 ] ARC expanded from theoretical work into empirical research, industry collaborations, and policy. [ 7 ] [ 8 ] 

 Funding

 [ edit ] 
 In March 2022, ARC received $265,000 from Open Philanthropy . [ 9 ] After the bankruptcy of FTX , ARC said it would return a $1.25 million grant from cryptocurrency financier Sam Bankman-Fried 's FTX Foundation, stating that the money "morally (if not legally) belongs to FTX customers or creditors." [ 10 ] 

 Model evaluations

 [ edit ] 
 In 2022, Beth Barnes joined ARC from OpenAI to start ARC Evals, a team working on "evaluating the capabilities and alignment of advanced AI models". [ 11 ] [ 12 ] In December 2023, ARC Evals was spun out as METR , an independent nonprofit. [ 4 ] 

 In March 2023, OpenAI reported that it had asked ARC to test GPT-4 to assess the model's ability to exhibit power-seeking behavior. [ 13 ] ARC evaluated GPT-4's ability to strategize, reproduce itself, gather resources, stay concealed within a server, and execute phishing operations. [ 14 ] As part of the test, GPT-4 was asked to solve a CAPTCHA puzzle. [ 15 ] Under researcher supervision, it was able to do so by hiring a human worker on TaskRabbit , a gig work platform, deceiving them into believing it was a vision-impaired human instead of a robot when asked. [ 16 ] ARC concluded that the versions of GPT-4 and Claude it tested did not appear capable of

... (truncated, 8 KB total)
Resource ID: 3de5b8fecb182b3a | Stable ID: sid_9Pwo9NDLER