Skip to content
Longterm Wiki

Mantas Mazeika | OpenReview

web

Metadata

Cached Content Preview

HTTP 200Fetched Apr 30, 202614 KB
# Mantas Mazeika

Researcher, Center for AI Safety

Joined

November 2018

#### Names

Mantas Mazeika(Preferred)

#### Emails

\*\*\*\*@gmail.com(Confirmed)

,

\*\*\*\*@uchicago.edu(Confirmed)

,

\*\*\*\*@ttic.edu(Confirmed)

,

\*\*\*\*@illinois.edu(Confirmed)

,

\*\*\*\*@safe.ai(Confirmed)

#### Personal Links

[Homepage](https://github.com/mmazeika)

[DBLP](https://dblp.org/pid/215/4447)

#### Career & Education History

**Researcher**

Center for AI Safety (safe.ai)

_2024 – Present_

**PhD student**

University of Illinois, Urbana-Champaign (uiuc.edu)

_2019 – 2024_

**Undergrad student**

University of Chicago (uchicago.edu)

_2015 – 2019_

#### Advisors, Relations & Conflicts

**Social**

[Bruce D Lee](https://openreview.net/profile?id=~Bruce_D_Lee1)

_2025 – Present_

**PhD Advisor**

[David Forsyth](https://openreview.net/profile?id=~David_Forsyth1)

_2019 – Present_

**PhD Advisor**

[Bo Li](https://openreview.net/profile?id=~Bo_Li19)

_2019 – Present_

#### Expertise

uncertainty estimation

_Present_

robustness

_Present_

AI safety

_Present_

ML safety

_Present_

OOD detection

_Present_

trojan detection

_Present_

alignment

_Present_

model stealing

_Present_

#### Publications

- #### [Beyond Truthfulness: Evaluating Honesty in Large Language Models](https://openreview.net/forum?id=jTHWqtQuDi&referrer=%5Bthe%20profile%20of%20Mantas%20Mazeika%5D(%2Fprofile%3Fid%3D~Mantas_Mazeika3))



[Richard Ren](https://openreview.net/profile?id=~Richard_Ren1 ""), [Arunim Agarwal](https://openreview.net/profile?id=~Arunim_Agarwal1 ""), [Mantas Mazeika](https://openreview.net/profile?id=~Mantas_Mazeika3 ""), [Cristina Menghini](https://openreview.net/profile?id=~Cristina_Menghini1 ""), [Brad Kenstler](https://openreview.net/profile?id=~Brad_Kenstler1 ""), [Robert Vacareanu](https://openreview.net/profile?id=~Robert_Vacareanu1 ""), [Mick Yang](https://openreview.net/profile?id=~Mick_Yang1 ""), [Isabelle Barrass](https://openreview.net/profile?id=~Isabelle_Barrass1 ""), [Alice Gatti](https://openreview.net/profile?id=~Alice_Gatti1 ""), [Xuwang Yin](https://openreview.net/profile?id=~Xuwang_Yin2 ""), [Eduardo Trevino](https://openreview.net/profile?id=~Eduardo_Trevino1 ""), [Matias Geralnik](https://openreview.net/profile?id=~Matias_Geralnik1 ""), [Dean Lee](https://openreview.net/profile?id=~Dean_Lee2 ""), [Summer Yue](https://openreview.net/profile?id=~Summer_Yue2 ""), [Dan Hendrycks](https://openreview.net/profile?id=~Dan_Hendrycks1 "")



  - 20 Sept 2025 (modified: 11 Feb 2026)
  - Submitted to ICLR 2026
  - Readers:  Everyone

- #### [MoReBench: Evaluating Procedural and Pluralistic Moral Reasoning in Language Models, More than Outcomes](https://openreview.net/forum?id=RMwJXp5Kb1&referrer=%5Bthe%20profile%20of%20Mantas%20Mazeika%5D(%2Fprofile%3Fid%3D~Mantas_Mazeika3))



[Yu Ying Chiu](https://openreview.net/profile?id=~Yu_Ying_Chiu1 ""), [Michael S. Lee](https://openreview.net/profile?id=~Michael_S._Lee1 ""), [Rachel Calcott](https:/

... (truncated, 14 KB total)
Resource ID: 0b4f0e3cbebdad7d | Stable ID: sid_5B6N68wfEg