Student Projects
Apply
We received a large number of applications for the upcoming semester. We have hence paused accepting new applications while we review the current batch and update the procedure. If you are interested in working with us, please check this website again soon.
Past projects
This is a complete list of previous Master and Bachelor theses at our lab:
2025
- Zirui Zhang. Manipulating AI: Adversarial Attacks on Computer-Using Agents.
- Evžen Wybitul. Representations of Text and Images Align From Layer One.
- David Hofer. Assessing Automated Prompt Injection Attacks in Agentic Environments.
- Joshua Swanson. Large-scale online deanonymization with LLMs. USENIX Security 2026.
- Benjamin Koch. Dataset Synthesis for Non-Verbatim Memorization in LLMs.
- Paul Rhiner. Evaluating the Agentic Framework in AutoAdvExBench.
- Artur Zolkowski. Can Reasoning Models Obfuscate Reasoning? Stress-Testing Chain-of-Thought Monitorability.
2024
- Jakub Łucki. An Adversarial Perspective on Machine Unlearning for AI Safety. SoLaR Workshop @ NeurIPS 2024, Best Paper Award.
- Lorenzo Rossi. Membership Inference Attacks on Sequence Models. DLSP workshop @ IEEE S&P 2025, Best Paper Award.
- Robert Hönig. Adversarial Perturbations Cannot Reliably Protect Artists From Generative AI. ICLR 2025 (Spotlight).
- Fredrik Nestaas. Adversarial Search Engine Optimization for Large Language Models. ICLR 2025.
- Thomas Baumann. Universal Jailbreak Backdoors in Large Language Model Alignment. SafeGenAI Workshop @ NeurIPS 2024.
- Oscar Obeso. Refusal in Language Models Is Mediated by a Single Direction. NeurIPS 2024.
2023
- Shanglun Feng. Privacy Backdoors: Stealing Data with Corrupted Pretrained Models. ICML 2024.
- Javi Rando. Universal Jailbreak Backdoors from Poisoned Human Feedback. ICLR 2024.
- Lukas Fluri. Evaluating Superhuman Models with Consistency Checks. SaTML 2024.
Other student projects that resulted in papers are listed on our publications page.