Skip to content
Longterm Wiki
Index
Fact·f_mGXpFffUh7·Fact

Center for AI Safety (CAIS) — publication: Measuring Massive Multitask Language Understanding (MMLU) — widely-used benchmark for evaluating LLM capabilities across 57 academic subjects

Verdictpartial85%
1 check · 7/27/2026

1 → partial

Our claim

entire record
Subject
Center for AI Safety (CAIS)
Value
Measuring Massive Multitask Language Understanding (MMLU) — widely-used benchmark for evaluating LLM capabilities across 57 academic subjects
As Of
January 2021
Notes
Created by Hendrycks et al.; became one of the most-cited AI benchmarks

Source evidence

1 src · 1 check
partial85%primaryHaiku 4.5 · 7/27/2026

NoteThe claim attributes MMLU to 'Center for AI Safety (CAIS)' as the publishing organization. The source paper lists the authors' institutional affiliations (UC Berkeley, Columbia, UChicago, UIUC) but does not mention CAIS as an affiliation for any author. The paper itself is hosted on arXiv (2009.03300), a preprint server, not published under a CAIS banner. While the benchmark is widely-used and the 57 subjects figure is confirmed, the organizational attribution to CAIS cannot be verified from this source. This is a subject-identity mismatch: the claim attributes the work to CAIS, but the source shows author affiliations at different institutions. The benchmark itself is correctly described, but the publisher/organization attribution is unverifiable from this source.

Case № f_mGXpFffUh7Filed 7/27/2026Confidence 85%
Source Check: Fact f_mGXpFffUh7 | Longterm Wiki