Center for AI Safety (CAIS) — Description: The Center for AI Safety (CAIS) is a San Francisco-based nonprofit focused on reducing societal-scale risks from AI through technical safety research, field-building, and public communication. Founded by Dan Hendrycks and Oliver Zhang. Known for the MMLU benchmark, representation engineering, and the May 2023 "Statement on AI Risk" signed by 350+ AI leaders.
1 → partial
Our claim
entire record- Subject
- Center for AI Safety (CAIS)
- Property
- Description
- Value
The Center for AI Safety (CAIS) is a San Francisco-based nonprofit focused on reducing societal-scale risks from AI through technical safety research, field-building, and public communication. Founded by Dan Hendrycks and Oliver Zhang. Known for the MMLU benchmark, representation… expand
The Center for AI Safety (CAIS) is a San Francisco-based nonprofit focused on reducing societal-scale risks from AI through technical safety research, field-building, and public communication. Founded by Dan Hendrycks and Oliver Zhang. Known for the MMLU benchmark, representation engineering, and the May 2023 "Statement on AI Risk" signed by 350+ AI leaders.- As Of
- 2025
Source evidence
1 src · 1 checkNoteThe claim makes four main assertions: (1) San Francisco-based nonprofit—unverifiable from source (location not stated); (2) focused on reducing societal-scale risks through technical safety research, field-building, and public communication—CONFIRMED (source says 'research, field-building, and advocacy'); (3) Founded by Dan Hendrycks and Oliver Zhang—CONFIRMED (both listed in Leadership); (4) Known for MMLU benchmark, representation engineering, and May 2023 'Statement on AI Risk' signed by 350+ leaders—CONTRADICTED on the statement signature count. The source states '600 leading AI researchers and public figures' signed a 'Global Statement on AI Risk,' not 350+. The MMLU benchmark and representation engineering are not mentioned in the source, making those claims unverifiable. The statement signature count is a direct contradiction (600 vs. 350+), which downgrades the overall verdict from 'confirmed' to 'partial.'