Mantas Mazeika | OpenReview
webMetadata
Cached Content Preview
HTTP 200Fetched Apr 30, 202614 KB
# Mantas Mazeika
Researcher, Center for AI Safety
Joined
November 2018
#### Names
Mantas Mazeika(Preferred)
#### Emails
\*\*\*\*@gmail.com(Confirmed)
,
\*\*\*\*@uchicago.edu(Confirmed)
,
\*\*\*\*@ttic.edu(Confirmed)
,
\*\*\*\*@illinois.edu(Confirmed)
,
\*\*\*\*@safe.ai(Confirmed)
#### Personal Links
[Homepage](https://github.com/mmazeika)
[DBLP](https://dblp.org/pid/215/4447)
#### Career & Education History
**Researcher**
Center for AI Safety (safe.ai)
_2024 – Present_
**PhD student**
University of Illinois, Urbana-Champaign (uiuc.edu)
_2019 – 2024_
**Undergrad student**
University of Chicago (uchicago.edu)
_2015 – 2019_
#### Advisors, Relations & Conflicts
**Social**
[Bruce D Lee](https://openreview.net/profile?id=~Bruce_D_Lee1)
_2025 – Present_
**PhD Advisor**
[David Forsyth](https://openreview.net/profile?id=~David_Forsyth1)
_2019 – Present_
**PhD Advisor**
[Bo Li](https://openreview.net/profile?id=~Bo_Li19)
_2019 – Present_
#### Expertise
uncertainty estimation
_Present_
robustness
_Present_
AI safety
_Present_
ML safety
_Present_
OOD detection
_Present_
trojan detection
_Present_
alignment
_Present_
model stealing
_Present_
#### Publications
- #### [Beyond Truthfulness: Evaluating Honesty in Large Language Models](https://openreview.net/forum?id=jTHWqtQuDi&referrer=%5Bthe%20profile%20of%20Mantas%20Mazeika%5D(%2Fprofile%3Fid%3D~Mantas_Mazeika3))
[Richard Ren](https://openreview.net/profile?id=~Richard_Ren1 ""), [Arunim Agarwal](https://openreview.net/profile?id=~Arunim_Agarwal1 ""), [Mantas Mazeika](https://openreview.net/profile?id=~Mantas_Mazeika3 ""), [Cristina Menghini](https://openreview.net/profile?id=~Cristina_Menghini1 ""), [Brad Kenstler](https://openreview.net/profile?id=~Brad_Kenstler1 ""), [Robert Vacareanu](https://openreview.net/profile?id=~Robert_Vacareanu1 ""), [Mick Yang](https://openreview.net/profile?id=~Mick_Yang1 ""), [Isabelle Barrass](https://openreview.net/profile?id=~Isabelle_Barrass1 ""), [Alice Gatti](https://openreview.net/profile?id=~Alice_Gatti1 ""), [Xuwang Yin](https://openreview.net/profile?id=~Xuwang_Yin2 ""), [Eduardo Trevino](https://openreview.net/profile?id=~Eduardo_Trevino1 ""), [Matias Geralnik](https://openreview.net/profile?id=~Matias_Geralnik1 ""), [Dean Lee](https://openreview.net/profile?id=~Dean_Lee2 ""), [Summer Yue](https://openreview.net/profile?id=~Summer_Yue2 ""), [Dan Hendrycks](https://openreview.net/profile?id=~Dan_Hendrycks1 "")
- 20 Sept 2025 (modified: 11 Feb 2026)
- Submitted to ICLR 2026
- Readers: Everyone
- #### [MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes](https://openreview.net/forum?id=RMwJXp5Kb1&referrer=%5Bthe%20profile%20of%20Mantas%20Mazeika%5D(%2Fprofile%3Fid%3D~Mantas_Mazeika3))
[Yu Ying Chiu](https://openreview.net/profile?id=~Yu_Ying_Chiu1 ""), [Michael S. Lee](https://openreview.net/profile?id=~Michael_S._Lee1 ""), [Rachel Calcott](https:/
... (truncated, 14 KB total)Resource ID:
0b4f0e3cbebdad7d | Stable ID: sid_5B6N68wfEg