Home | Publications | Wan26

Evaluating and Mitigating Misalignment in Language Models

MCML Authors

Abstract

This dissertation advances the evaluation and mitigation of misalignment in large language models by addressing both flawed optimization signals and the challenge of representing diverse human values. It develops methods to detect reward hacking and analyze safety behavior, introduces targeted interventions to reduce unnecessary refusals without weakening safety, and proposes an efficient framework for preserving differing human perspectives during model alignment. (Shortened).

phdthesis Wan26


Dissertation

LMU München. Jun. 2026

Authors

X. Wang

Links

DOI

Research Area

 B2 | Natural Language Processing

BibTeXKey: Wan26

Back to Top