Skip to content
Longterm Wiki
Index
Fact·f_mGXpFffUh7·Fact

Center for AI Safety (CAIS) — publication: Measuring Massive Multitask Language Understanding (MMLU) — widely-used benchmark for evaluating LLM capabilities across 57 academic subjects

Verdictpartial85%
1 check · 9/21/2026

1 → partial

Our claim

entire record
Subject
Center for AI Safety (CAIS)
Value
Measuring Massive Multitask Language Understanding (MMLU) — widely-used benchmark for evaluating LLM capabilities across 57 academic subjects
As Of
January 2021
Notes
Created by Hendrycks et al.; became one of the most-cited AI benchmarks

Source evidence

1 src · 1 check
partial85%primaryHaiku 4.5 · 9/21/2026

NoteThe claim attributes MMLU to 'Center for AI Safety (CAIS)' as a publication. The source text is the arXiv paper itself (2009.03300) authored by Hendrycks et al. from UC Berkeley, Columbia, UChicago, UIUC, etc. The source makes no mention of CAIS as the publishing or affiliated organization. The paper's author affiliations are listed as academic institutions, not CAIS. While the benchmark details (57 subjects, Hendrycks et al. authorship, academic scope) are confirmed, the organizational attribution to CAIS is not addressed in the source. This is a subject-identity mismatch: the claim attributes the work to CAIS, but the source shows it as an independent academic paper without CAIS affiliation. The temporal marker 'as of 2021-01' is also not verifiable from the source (paper is from 2020-09).

Case № f_mGXpFffUh7Filed 9/21/2026Confidence 85%
Source Check: Fact f_mGXpFffUh7 | Longterm Wiki