Publications

Don’t Forget! Decomposing the Training Dynamics of Memorization in Language Models

Florian Eichin, Philipp Mondorf, Andrei Mircea, Yupei Du, Barbara Plank, Michael A. Hedderich. 2026. Don't Forget! Decomposing the Training Dynamics of Memorization in Language Models. Preprint.

We decompose LLM loss trajectories to characterize memorization training dynamics through gradient alignment and validate our results through early memorization prediction causal post-hoc ablation.

Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration

Ryan Soh-Eun Shim, Kwanghee Choi, Kalvin Chang, Ming-Hao Hsu, Florian Eichin, Zhizheng Wu, Alane Suhr, Michael A. Hedderich, David Harwath, David R. Mortensen, Barbara Plank. 2026. Linear Script Representations in Speech Foundation Models Enable Zero-Shot Transliteration. ACL 2026 Findings.

We show that script is largely represented as a single linear direction in activation space of Whisper models and that steering activations towards that direction enables transcriptions and even generalizes across different languages.